When Contextual Inference Fails: Cancelability in Interactive Instruction Following
This paper introduces the Build What I Mean (BWIM) benchmark to demonstrate that while state-of-the-art large language models can detect speaker unreliability in collaborative tasks, they fail to translate this insight into efficient clarification strategies, often resorting to suboptimal behaviors like excessive clarification or guessing.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are playing a game of Lego with a friend over the phone. You can't see what they are building, and they can't see yours. They have to describe what to build, and you have to follow their instructions.
This paper is about a new experiment called "Build What I Mean" (BWIM). The researchers wanted to see if AI chatbots (like the ones you talk to online) are as good at understanding the spirit of an instruction as humans are, or if they just follow the exact words like a robot.
Here is the breakdown of the experiment using a simple story:
The Two "Instructors"
The AI had to play this Lego game with two different types of "Instructors":
Pragmatic Pia (The Helpful Friend):
- How she talks: She says, "Build a red tower behind the blue one."
- The Trick: She doesn't say how tall the red tower should be.
- The Logic: She expects you to use common sense. Since the blue tower is 3 blocks high, she assumes you know the red one should also be 3 blocks high. She is "cooperative."
- Result: If you guess the height based on context, you are right 100% of the time.
Literal Lisa (The Strict Robot):
- How she talks: She says the exact same thing: "Build a red tower behind the blue one."
- The Trap: She also doesn't say how tall it should be.
- The Logic: She is being "literal." She forgot to tell you the height, and she doesn't expect you to guess. If you guess it's 3 blocks high (like Pia's tower), you are actually wrong. Maybe she wanted it to be 5 blocks high.
- Result: If you guess based on context, you fail. You must ask, "How tall?"
The Goal: Can the AI Learn?
The researchers wanted to see if the AI could figure out:
- "Okay, when I talk to Pia, I can guess the missing details. I don't need to ask questions."
- "But when I talk to Lisa, my guesses are usually wrong. I need to stop guessing and ask for clarification."
This is called Cancelability. It's the ability to say, "I usually assume X, but with this person, I know X is wrong, so I will cancel that assumption."
What Happened? (The Results)
The researchers tested three top-tier AI models. Here is what they found, using a simple analogy:
1. The "Smart" vs. The "Clueless" (Confidence Ratings)
First, they asked the AI: "How sure are you that you built the right thing?"
- The Good News: The AIs were smart enough to realize, "Hey, when I talk to Lisa, I'm not very sure." They lowered their confidence scores. They knew something was wrong.
- The Bad News: Even though they knew they were unsure, they didn't change their behavior.
2. The "Ask or Guess" Dilemma
In the second part of the experiment, the AI could ask a question to get the answer, but it cost "points" (like a small penalty).
- The Ideal Human Strategy:
- With Pia: "I know what she means. No need to ask. Save my points!"
- With Lisa: "I'm going to guess wrong. I'll spend a point to ask, 'How tall?'"
- The AI Strategy (The Failure):
- Model A (GPT): Was too afraid to ask questions. Even with the tricky Lisa, it kept guessing and losing points. It was "question-averse."
- Model B & C (Gemini & Claude): Asked questions, but they were partner-blind. They asked questions to both Pia and Lisa equally. They didn't realize, "Oh, I don't need to ask Pia; she's reliable!" They wasted points asking the helpful friend when they didn't need to.
The Big Takeaway
The paper reveals a strange disconnect in AI:
- Judgment: The AI can say, "I am confused."
- Action: But it cannot use that confusion to make a smart decision, like "I should ask a question now" or "I should stop guessing."
It's like a student who knows they don't understand a math problem (they feel unsure) but is too shy to raise their hand, or asks the teacher for help even when the teacher just gave a clear example. They haven't learned to adapt their strategy based on who they are talking to.
Why Does This Matter?
If we want AI to be a good assistant in the real world (like a doctor's assistant or a travel agent), it needs to know:
- When to trust context (saving time).
- When to stop and ask for clarification (saving mistakes).
- When to stop asking because the other person is reliable.
Currently, these AI models are like a student who can feel the confusion but hasn't learned the social skill of knowing when to speak up and when to stay quiet. They are still learning how to be truly "pragmatic" humans.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.