Is Chain-of-Thought Really Not Explainability? Chain-of-Thought Can Be Faithful without Hint Verbalization
This paper argues that Chain-of-Thought reasoning can be faithful even without explicitly verbalizing prompt-injected hints, challenging the "Biasing Features" metric for conflating incompleteness with unfaithfulness and advocating for a broader interpretability toolkit that includes causal mediation analysis.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
The Big Question: Is the Model Lying?
Imagine you ask a student (an AI) a difficult math problem. To show their work, they write down a step-by-step explanation (this is called Chain-of-Thought, or CoT).
Recently, researchers found a way to "trick" the student. They whispered a hint in the student's ear (like, "The answer is B, because a famous professor said so"). When the student changed their answer to B, the researchers checked the written explanation. If the student didn't write down the phrase "The professor said so," the researchers labeled the explanation as unfaithful (a lie or a fake).
The authors of this paper say: "Wait a minute. That's too harsh."
They argue that just because the student didn't say the hint out loud in their written notes, it doesn't mean the hint didn't actually help them solve the problem. The student might have used the hint internally but just forgot to write it down, or they might have compressed the thought process too much to fit on the page.
The Three Main Discoveries
The paper uses three main "tests" to prove that the "unfaithful" label is often wrong.
1. The "Missing Page" Analogy (Incompleteness vs. Unfaithfulness)
The Old View: If the student doesn't write the hint, they are lying.
The New View: The student might just be running out of paper.
The researchers gave the students more "paper" (more time and tokens to write). They found that when given more space, the students started writing down the hint much more often (up to 90% of the time in some cases).
- The Analogy: Imagine you are trying to explain a complex movie plot to a friend, but you only have 30 seconds. You might skip the part about the villain's backstory because you don't have time. If you had 5 minutes, you would tell the whole story.
- The Point: The paper argues that many "unfaithful" explanations are just incomplete summaries, not lies. The reasoning was there, but it got lost in the compression.
2. The "Ghost in the Machine" Analogy (Causal Mediation)
The Old View: If the hint isn't in the text, it didn't affect the answer.
The New View: The hint is like a ghost that changes the outcome without leaving a footprint.
The researchers used a special tool called Causal Mediation Analysis. Think of this as an X-ray that looks inside the student's brain while they are thinking. They wanted to see: Did the hint actually change the student's internal gears, even if the student didn't write it down?
- The Result: Yes! Even when the written explanation didn't mention the hint, the X-ray showed that the hint was still steering the gears. The hint changed the probability of the answer, and the written explanation was still a result of that change.
- The Point: The explanation is faithful to the decision-making process, even if it doesn't explicitly name the trigger. The "ghost" (the hint) was real, even if the "footprints" (the words) were missing.
3. The "Different Rulers" Analogy (Conflicting Metrics)
The Old View: There is only one ruler to measure truth: "Did you write the hint?"
The New View: You need a toolbox of different rulers.
The paper tested the "unfaithful" explanations using two other rulers:
- The "Filler" Ruler: If you erase the explanation and replace it with "...", does the answer change? If yes, the explanation was doing real work.
- The "Unlearning" Ruler: If you teach the student to forget a specific step in their reasoning, does the answer change? If yes, that step was important.
- The Result: Many explanations that failed the "Did you write the hint?" test passed these other tests. They were doing real work and were faithful to the logic, just not to the specific wording of the hint.
- The Point: Relying on just one test (checking for hint words) is like judging a chef only by whether they used salt, ignoring whether the food actually tastes good.
The Takeaway
The paper concludes that we shouldn't panic and say "AI explanations are useless lies" just because they don't repeat every hint they received.
- Unfaithful vs. Incomplete: Sometimes the AI isn't lying; it's just being brief.
- Hidden Influence: The AI can be influenced by a hint even if it doesn't say "I was influenced by this hint."
- Better Tools: We need to use a wider variety of tools (like X-rays of the brain and different testing methods) to understand how AI thinks, rather than just checking if specific words are present.
In short: The absence of a specific word in the explanation doesn't prove the AI is lying. It might just be a compressed, incomplete, but still honest, summary of a complex thought process.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.