← Latest papers
🤖 machine learning

When Chain-of-Thought Fails, the Solution Hides in the Hidden States

By using activation patching to transfer hidden states from reasoning traces to direct-answer runs, the authors demonstrate that individual Chain-of-Thought tokens encode sufficient task-relevant information to recover correct answers even when the original reasoning trace is flawed.

Original authors: Houman Mehrafarin, Amit Parekh, Ioannis Konstas

Published 2026-04-28
📖 4 min read☕ Coffee break read

Original authors: Houman Mehrafarin, Amit Parekh, Ioannis Konstas

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are watching a student take a difficult math exam. They are "thinking out loud," writing down every single step on their scratchpad. Suddenly, they reach the final answer and write: "The answer is 12."

You look at their work, do the math yourself, and realize they made a mistake. The real answer is 18.

Most people would assume that because the student wrote "12," their brain must have been totally lost. But what if I told you that if you could peek into their brain at the exact moment they were writing a specific word—like the word "subtract"—you would see that they actually knew the right answer, but they just stumbled at the very end?

That is exactly what this research paper is about.

The Core Idea: The "Hidden Genius" in the Mistake

The researchers wanted to know: When an AI (like ChatGPT) gives a wrong answer after explaining its reasoning (Chain-of-Thought), is the AI actually "stupid" during the whole process, or is the correct information just "hiding" inside its internal thoughts?

To find out, they used a technique called "Activation Patching."

The Analogy: The Radio Signal
Think of the AI's reasoning like a radio broadcast.

  • The "CoT" (Chain-of-Thought) is the actual song playing out loud (the text you read).
  • The "Hidden States" are the invisible radio waves traveling through the air.

Sometimes, the "song" (the text) is a distorted, messy version of the music. But the researchers discovered that even when the song sounds terrible and wrong, the "radio waves" (the internal math) are often still broadcasting a perfectly clear, correct signal.

By "patching"—which is like grabbing a clean signal from one moment and plugging it into another—they were able to "fix" the AI's broken reasoning and force it to output the correct answer, even when the AI's own written explanation was wrong.

The Three Big Discoveries

1. The "Verb" vs. "Number" Secret

The researchers found that not all words are created equal.

  • The "Logic" Words (Verbs and Entities): Words like "subtract," "total," or "Janet" are like the GPS coordinates of the problem. If you "patch" these into the AI, it suddenly finds its way to the right answer. These words carry the strategy.
  • The "Math" Words (Numbers and Operators): Words like "16" or "+" are like the passengers in the car. They are just numbers floating around. If you patch just a number, the AI often just guesses a random wrong number. It doesn't have the "map" to get home.

2. The "Early Bird" Advantage

The researchers found that the most useful "correct" information usually appears early in the reasoning process.
The Analogy: It’s like a chef preparing a meal. If they pick the wrong ingredients at the very beginning, the meal is ruined. But if they have the right ingredients (the correct logic) in their hands early on, they can still cook a perfect meal even if they get a bit messy during the actual cooking.

3. You Don't Need a Long Story

Usually, we think AI needs to write a long, beautiful essay to solve a hard problem. But this paper shows that the AI's "internal brain" is much more efficient than its "mouth."
The researchers found they could "patch" a single internal thought into the AI, and it would spit out the correct answer in just a few words. It didn't need the long, rambling explanation to get it right; it just needed that one "spark" of correct internal logic.

Why does this matter?

This research tells us that AI "reasoning" is deeper than the words it types.

When an AI fails, it isn't always because it doesn't know the answer; it's often because it fails to translate its internal knowledge into the written words we see. This opens the door to building smarter AIs that don't need to write long, rambling paragraphs to be accurate—they just need to keep their "internal radio signal" clear.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →