NewsInterview: a Dataset and a Playground to Evaluate LLMs' Ground Gap via Informational Interviews
This paper introduces NewsInterview, a dataset of 40,000 informational interviews and a simulated environment that reveals Large Language Models' significant deficits in strategic dialogue, such as recognizing answered questions and pivoting effectively, compared to human interviewers.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are at a dinner party. You have a friend who is a master storyteller, but they are also a bit shy, nervous, or maybe even a little grumpy. Your goal is to get them to tell you their most fascinating, secret stories.
To do this well, you need more than just a list of questions. You need to:
- Listen and nod ("Oh wow, that sounds scary!") to make them feel safe.
- Read the room to know when to push for more details and when to change the subject.
- Build trust so they eventually open up completely.
This paper, titled "NewsInterview," is about testing whether Artificial Intelligence (specifically Large Language Models or LLMs) can play the role of that skilled dinner party host. The researchers found that while AI is great at writing essays, it is surprisingly bad at having a strategic conversation.
Here is the breakdown of their discovery, using some fun analogies:
1. The Problem: The "Robot Interviewer"
The researchers looked at 40,000 real interviews from NPR and CNN. They found that human journalists are like skilled gardeners. They know exactly when to water a plant (ask a follow-up), when to prune a branch (change the topic), and when to just sit quietly and smile (acknowledge the guest's feelings).
When they asked an AI to act as the interviewer, the AI behaved more like a robotic vending machine.
- No "Nods": Humans use "acknowledgment statements" (like "I see," or "That's interesting") about 9% of the time to build trust. The AI almost never did this. It was like talking to someone who never says "uh-huh" or "wow."
- The Rabbit Hole: Instead of moving the conversation forward to new topics (like a gardener moving to the next flower bed), the AI kept asking follow-up questions on the same topic, getting stuck in a loop.
- Missing the Strategy: Humans have a plan. They know they need to get three specific facts out of the guest. The AI just asked questions randomly, hoping something good would happen.
2. The Solution: The "Interview Video Game"
To fix this, the researchers built a simulated video game called NewsInterview.
- The Setup: One AI plays the Interviewer, and another AI plays the Source (the guest).
- The Characters: The "Source" AI isn't just one person; it wears different masks (personas). Sometimes it's Anxious (scared to talk), Defensive (angry), Clueless (doesn't know the answer), or Dominating (won't stop talking).
- The Rules: The Interviewer AI gets points only if it successfully "persuades" the Source to reveal hidden information.
- Analogy: Imagine the Source is a locked safe. The Interviewer has to pick the lock. If the safe is "Anxious," you have to whisper and be gentle. If the safe is "Adversarial," you have to be firm and show your ID. If you just bang on the safe with a hammer (ask blunt questions), it won't open.
3. The Results: Who Won the Game?
The researchers ran the game thousands of times. Here is what happened:
- The "Source" AI was great: When the AI played the role of the guest, it was very good at acting human. If it was told to be "Anxious," it acted nervous. If it was told to be "Defensive," it got grumpy. It even knew when it was being persuaded.
- The "Interviewer" AI struggled: Even the smartest AI models (like GPT-4) failed to extract the maximum amount of information.
- They couldn't tell when a question had been answered and when to move on.
- They couldn't adapt their style. If the guest was scared, the AI didn't know to be gentle. It just kept asking the same tough questions, which made the "guest" shut down.
- The Score: In the hardest version of the game (where the guest is tricky), the AI only managed to get about 50% of the available information. When the game was made easier (removing the "persuasion" part), the score jumped to over 80%. This proves the AI's main weakness isn't knowing facts; it's knowing how to talk to people.
4. Why Does This Matter?
You might think, "So what? It's just a news interview." But the authors argue that strategic conversation is the key to many real-world jobs:
- Therapists: Need to build trust to help patients open up.
- Teachers: Need to encourage students to keep trying.
- Crisis Negotiators: Need to talk down a hostage-taker.
Currently, AI is like a student who has read every textbook but has never actually talked to a human. It knows the words, but it doesn't understand the dance of conversation.
The Takeaway
This paper gives us a new playground to teach AI how to be a better conversationalist. It shows that to make AI truly helpful in sensitive situations, we can't just teach it more facts. We have to teach it empathy, timing, and the art of persuasion—the same skills a human journalist uses to get the story.
In short: AI is a brilliant librarian, but it's currently a terrible detective. This paper is the training manual to teach it how to solve the case.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.