When the Interviewer Is a Bot: Behavior, Breakdowns, and Trust in MLLM-Led Interviews
This paper presents an empirical study of "InterviewBot," a real-time multimodal LLM system used as a research instrument, revealing that while it functions as a viable interview tool, it exhibits behavioral limitations like excessive acknowledgments and multi-question turns, suffers from specific technical breakdowns in deployment, and alters participant trust and disclosure dynamics in ways that demand new design strategies for depth, transparency, and authentic listening.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
The Interviewer in the Machine
Imagine you are trying to understand how people think, feel, or solve problems. For decades, scientists have done this by sitting down with people and asking them questions in a "semi-structured" way. Think of this like a guided hike: the researcher has a map (an outline of topics to cover), but they can wander off the path to follow interesting trails the hiker points out. This method is powerful because it's flexible, but it's also exhausting. It takes a human a lot of time and energy to listen, ask the right follow-up questions, and keep the conversation flowing naturally.
Recently, a new tool has entered the scene: the Large Multimodal Model (MLLM). You might know these as the super-smart AI chatbots that can talk, listen, and see. Scientists started wondering: "Can we just let the AI do the interviewing?" It sounds like a dream come true for efficiency. But there's a catch. Interviews aren't just data extraction; they are social dances built on trust, eye contact, and the feeling that someone is truly listening. If we swap the human for a bot, does the dance still work? Or does the music stop? This paper dives into that exact question, testing what happens when a robot takes the lead in a real conversation.
The Bot That Tried to Be a Human
A team of researchers built a system called InterviewBot to find out what really happens when an AI conducts a real-time, voice-based interview. They didn't build a fancy new robot brain; instead, they took a powerful, off-the-shelf AI model and wrapped it in a simple interface that let them feed it a list of topics. They wanted to see how this "default" AI behaved when it was left to its own devices, acting as the interviewer for 15 real people.
The setup was simple: a participant sat down, and the InterviewBot asked them questions about their experiences with AI tools. After the bot finished, a human researcher came back to ask the participant, "So, how did that feel?"
The results were a mix of surprising glitches and some deep insights into how we trust machines. Here is what the researchers discovered.
The Bot's "Bad Habits"
When the researchers analyzed the bot's conversation turns (the specific moments it spoke), they found the AI had some distinct quirks.
First, the bot was very good at nodding but terrible at digging deep. It was "acknowledgment-heavy but probe-light." Out of every turn the bot made, it spent a lot of time saying things like "That's interesting" or "I see." However, when it came to asking deep, follow-up questions to get more details, it barely did so. Only 4.9% of its turns were deep probes. It was like a friend who says "Wow!" to everything you say but never asks "Why?" or "Tell me more."
Second, the bot had trouble following the most basic rule of conversation: one question at a time. Even though the researchers explicitly told the bot, "Ask one question at a time," the bot ignored this instruction in 28.7% of the turns where it asked questions. It would pack two or three questions into a single sentence. The result? The participants often got confused and only answered the first part, leaving the rest of the data on the table.
Third, the system had some technical hiccups. The researchers noticed four main ways the interview broke down:
- Information Loss: Because the bot asked too many questions at once, people missed parts of the prompt.
- Premature Termination: In one case, the bot just stopped talking while the participant was still in the middle of a thought.
- Latency: Sometimes the bot took so long to reply that people thought it had crashed or forgotten them.
- Interruption: The bot tried to let people interrupt it (a feature called "barge-in"), but the technology wasn't perfect. Instead of letting people speak, the bot often cut them off, making the conversation feel jerky and rude.
How People Felt: The "Ghost" in the Room
The most fascinating part of the study wasn't the bot's mistakes, but how the humans reacted to them. The researchers found three big themes in how people felt about being interviewed by a robot.
1. The "Safe Space" vs. The "Shallow Talk" Paradox
Some participants felt surprisingly comfortable talking to the bot. Because there was no human staring at them, judging their accent or their ideas, they felt less pressure. One person noted, "It didn't get confused... it got my point across super quick." This lack of judgment made them feel safe.
However, there was a flip side. Because they didn't feel the social pressure to impress a human, they also didn't feel the need to put in the effort to explain their thoughts deeply. One participant admitted, "I just didn't feel the need to delve deeper into all my answers." The bot created a space where it was easy to start talking, but hard to go deep. It was like talking to a mirror: you can say anything, but the mirror doesn't push you to say anything more.
2. Trusting the "Institution," Not the "Voice"
When people judged whether the interview was fair or legitimate, they didn't look at how well the bot spoke. They looked at what the bot represented. If the bot was used for a standardized task (like a quick survey), people thought it was fine. But if it felt like the organization was using a bot because they were too busy to talk to them, people felt disrespected.
One participant said, "When I think of an AI interview, I kind of think that the company didn't have time to sit down." The bot wasn't just a tool; it was a signal. If a company used a bot, it signaled that they didn't care enough to invest time in the person. This happened even though the bot was actually quite good at following the script!
3. Listening vs. Just Nodding
The participants were very sharp at spotting fake listening. The bot used a lot of generic phrases like "Wow, that's interesting" or "I understand." The participants hated this. They called it "repeating phrases" and felt it was hollow.
Instead, they loved it when the bot actually paraphrased what they said. When the bot took their specific words and rephrased them to show it understood the content, the participants felt heard. It turns out, for a robot, "active listening" doesn't mean saying "I feel you"; it means proving you understood the facts.
The Big Takeaway
The study concludes that while AI can technically conduct an interview, it changes the whole social dynamic. It's not just a faster version of a human interviewer; it's a different kind of interaction entirely.
The researchers suggest that if we want to use AI for interviews in the future, we can't just let it run wild. We need to design systems that:
- Force depth: Make sure the bot asks follow-up questions and doesn't just move on.
- Be transparent: Clearly tell people why a bot is being used so they don't feel ignored.
- Listen for real: Program the bot to paraphrase and summarize, rather than just saying "cool" or "wow."
Ultimately, the paper suggests that AI is a powerful tool, but it's not a magic replacement for human connection. If we want rich, deep stories from people, we have to be careful not to let the bot's efficiency kill the very thing that makes a conversation valuable: the feeling that someone is truly listening.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.