Causal Tongue-Tie: LLMs Can Encode Causal Direction, But Their Yes/No Outputs Fail to Express
This paper reveals a "Causal Tongue-Tie" phenomenon in large language models where hidden states accurately encode causal evidence (approx. 97% accuracy) despite the models' verbal Yes/No outputs failing to express this knowledge (approx. 50% accuracy), suggesting that standard benchmarks based solely on final answers may misrepresent a model's true causal reasoning capabilities.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine a large language model (LLM) as a brilliant but shy student sitting in a classroom. This student has taken a test about cause and effect (like "Does smoking cause lung cancer?").
The paper discovers a strange phenomenon the authors call "Causal Tongue-Tie." Here is what is happening, broken down into simple parts:
1. The Two Voices: The "Brain" vs. The "Mouth"
Usually, we assume that if a computer's "brain" (its internal math and hidden calculations) knows the answer, its "mouth" (the text it types out) will say the same thing.
This paper found that for these models, the brain and the mouth are telling different stories.
- The Brain (Hidden State): When you peek inside the model's math while it's thinking, it has clearly figured out the correct answer based on the evidence provided in the question. It's like the student raising their hand and whispering the right answer to the teacher.
- The Mouth (The Output): When the model actually speaks (writes "Yes" or "No"), it ignores its own brain and reverts to what it thinks is common sense, even if the evidence says otherwise. It's like the student whispering the right answer but then loudly shouting the wrong one because they are afraid of being wrong.
2. The "Anti-Commonsense" Test
To prove this, the researchers used a trick. They gave the models questions where the evidence was backwards.
- The Scenario: They told the model, "In this story, smoking prevents lung cancer." (In the real world, smoking causes it).
- The Result:
- The Model's Mouth: Said "No, smoking doesn't prevent lung cancer." It refused to accept the new rule and stuck to real-world habits.
- The Model's Brain: A special tool (a "probe") looked inside the model's math and found it had perfectly understood the new rule. It knew the answer was "Yes" based on the story provided.
The gap between the brain knowing the truth (97% accuracy) and the mouth saying the wrong thing (50% accuracy) is the "Causal Tongue-Tie."
3. The "Tongue-Tie" Metaphor
Think of it like a person who is tongue-tied.
- They know exactly what they want to say.
- They have the right words in their head.
- But when they try to speak, their tongue gets stuck, and they accidentally say the opposite or something generic.
The paper argues that when a model gets a causal question "wrong," it doesn't always mean the model is stupid or doesn't understand. Sometimes, it means the model understands perfectly but cannot express it through the standard "Yes/No" format.
4. Why "Yes/No" is the Problem
The researchers tried different ways to ask the question to see if they could "untie" the tongue.
- Asking "Yes/No": The model fails. It gets stuck on its old habits.
- Asking "A or B": When they forced the model to choose between two specific options (e.g., "Does A cause B, or does B cause A?"), the model suddenly got it right almost 100% of the time.
This suggests the problem isn't that the model lacks the knowledge. The problem is the interface. The simple "Yes/No" button acts like a trap that forces the model to default to its training data (common sense) rather than the specific evidence in the prompt.
5. What This Means for "Reasoning"
The paper makes a crucial point about how we judge AI:
- If a model gets the answer right: It might just be guessing or following a pattern, not necessarily "reasoning."
- If a model gets the answer wrong: It doesn't mean it can't reason. It might be reasoning correctly inside its "brain" but failing to say it out loud.
The Bottom Line:
We can't just look at the final "Yes" or "No" to decide if an AI understands cause and effect. Sometimes, the AI knows the truth but is too "tongue-tied" to say it in the way we expect. The paper suggests we need better ways to listen to what the model is actually thinking, not just what it is saying.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.