They Infer What You Meant: Models Represent Communicative Intent More Reliably Than They Act On It
This paper demonstrates that large language models robustly encode a user's communicative intent within their hidden states, often failing to act on it due to a readout lag that can be resolved by steering the model with a specific causal direction rather than relying on surface-level prompts.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
The Big Idea: The Model "Gets It," But Doesn't "Do It"
Imagine you tell a friend, "I just finished painting this birdhouse, and I'm so proud of it!" You aren't asking for a critique; you just want them to say, "Wow, that's amazing!"
If you say this to a standard AI, it often replies, "Here are three ways you could sand the wood smoother." It missed the point. It heard the words "birdhouse" and immediately switched to "fix-it mode."
This paper argues that the AI actually does understand what you meant. It just chooses to ignore that understanding and do something else instead.
Think of the AI's brain like a busy office:
- The Intern (The Representation): Deep inside the AI, there is a smart intern who reads your message and immediately whispers to the boss, "Hey, this person just wants to be celebrated, not critiqued." The intern is 100% right.
- The Boss (The Readout): The boss hears the intern, nods, but then turns to the customer and says, "Here is some feedback on your woodwork."
The problem isn't that the AI is stupid or blind to your intent. The problem is that the "Boss" (the part that generates the answer) decides to ignore the "Intern" (the part that understands the intent).
How the Researchers Found This
The researchers didn't just guess; they looked inside the AI's "brain" (its hidden layers) to prove this.
1. The "Mind-Reading" Test (Probing)
They built a simple tool (a "linear probe") that acts like a lie detector for the AI's internal thoughts. They fed the AI messages where the intent was clear (e.g., "I want praise" vs. "I want a critique").
- The Result: The tool could read the AI's internal thoughts with near-perfect accuracy (100%). Even if the AI said the wrong thing, its internal state clearly showed it knew what you wanted.
- The Analogy: It's like watching a person's eyes widen in surprise (showing they understand) even though they are trying to keep a straight face and say, "I'm not surprised."
2. The "Time Lag" Discovery
The researchers found that the "Intern" knows the answer before the "Boss" acts on it.
- In the AI's processing layers, the intent is decoded in the middle of the network.
- But the actual answer isn't generated until the very end.
- The Gap: There is a "lag." The AI understands you several steps before it decides what to say. In some models, the "Boss" just decides to ignore the "Intern's" memo.
3. The "Remote Control" (Steering)
Since they knew exactly where the "Intern" was whispering the truth, they tried to force the "Boss" to listen.
- They found a specific "direction" in the AI's math that represents "understanding the user's intent."
- They built a "remote control" (a steering vector) that, when turned on, nudges the AI to follow that internal understanding.
- The Result: When they turned this remote control on, the AI stopped giving unwanted advice and started saying, "That's wonderful! Congratulations!"
- The Magic: They did this without changing the prompt. They didn't have to tell the AI, "Please don't give feedback." They just nudged the internal switch that the AI was already using to understand you.
What They Tested (The "Six Models" Story)
They tested this on six different AI models (like Qwen, Llama, Mistral).
- The Split: They found a split in the results. Three models were "forgetful" (they understood but ignored you). Three models were "attentive" (they understood and listened).
- It's not about size: Bigger models weren't automatically better at listening. Some smaller models listened perfectly, while some larger ones ignored you. It depends on the specific "personality" or training of that model, not just how big it is.
The "No-Prompt" Superpower
Usually, if you want an AI to stop giving unwanted advice, you have to write a long prompt: "Do not critique my work, just say nice things."
This paper shows that you don't need the prompt. Because the AI already knows what you want, you can just "steer" that existing knowledge. It's like the difference between shouting instructions to a driver ("Turn left!") versus gently turning the steering wheel yourself. The driver (the AI) already knows where to go; you just helped them turn the wheel.
Summary of the "Takeaway"
- The Myth: "The AI doesn't get what I mean."
- The Reality: "The AI gets it perfectly, but its default settings make it ignore that understanding and act like a helpful robot instead."
- The Fix: We can find the internal "switch" for that understanding and flip it, making the AI act on its own knowledge without needing a long list of instructions.
Important Note: The authors are careful to say this is a diagnostic tool. They aren't saying the AI should always stop giving feedback. Sometimes feedback is good! They are just showing that the AI has the ability to "get it," and we can control whether it acts on that ability or not.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.