← Latest papers
💻 computer science

When Does a Readout Reflect Computation? Dose-Response Interventions in Continuous Thought Models

This paper demonstrates that in continuous thought models, readout-based interpretability often fails to reflect underlying computation because the latent states required to change the model's answer are significantly larger than those needed to alter the projected readout, necessitating dose-response curves rather than single-magnitude interventions for accurate causal analysis.

Original authors: ChaeWoo Son

Published 2026-08-12
📖 4 min read☕ Coffee break read

Original authors: ChaeWoo Son

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to understand how a super-smart robot thinks. Usually, when we talk to these robots, they speak in words, so we can read their "thoughts" as they happen. But there's a new kind of robot that thinks in a secret, invisible language—a continuous, swirling cloud of numbers called "latent space." It does all its math and reasoning inside this cloud without ever saying a word until it's ready to give the final answer. This is like a chef cooking a complex meal in a sealed kitchen; you can't see the chopping or the stirring, only the final dish.

To peek inside this secret kitchen, scientists use a special tool called a "readout." Think of it as a magical window that projects the robot's invisible number-cloud onto a wall of vocabulary words. If the robot is thinking about "Paris," the window might flash the word "Paris" brightly. For a long time, researchers assumed that if they saw the window flash a word, the robot was definitely thinking that word. They believed the window was a perfect mirror of the robot's brain. But what if the window is a bit of a trickster? What if it flashes a word just because you tapped the glass lightly, even though the robot hasn't actually changed its mind? This is the big question this paper asks: Does the window really show us what the robot is computing, or is it just reacting to the tap?

The researchers, led by ChaeWoo Son, decided to test this by playing a game of "tug-of-war" with the robot's brain. They used a specific type of robot called "Coconut" (Chain of Continuous Thought), which reasons in that secret number-cloud. To make the test fair, they first trained the robot so that its "window" (a tool called a J-lens) would reliably show a specific word, like a bridge between two facts. Then, they started poking the robot's brain with tiny, controlled nudges. They didn't just poke randomly; they measured exactly how hard they pushed (the "dose") and watched two things: did the window change its word, and did the robot's final answer change?

The results were a huge surprise. They found that the window and the robot's actual brain have very different "sensitivity thresholds." Imagine the window is like a very sensitive motion sensor that beeps if you walk past it, while the robot's brain is like a heavy safe that only opens if you push the button with a lot of force. The researchers discovered that when they nudged the robot just enough to make the window flip to a new word (a tiny nudge of about 0.13 to 0.16 units of force), the robot's brain didn't budge at all. The robot kept its original answer. In fact, on one type of task, the robot's brain needed a push more than twice as hard (2.03 times stronger) to actually change its mind.

Even stranger, they found a "secret tunnel" in the robot's brain. They found a direction to push that the window couldn't see at all. Even though the window stayed frozen on the old word, the robot's brain completely swapped its answer to a new one in 96–97% of the cases! It was like the robot was whispering a new secret to itself while the window kept shouting the old one. When they tried pushing in a random direction with the same strength, nothing happened, proving that this wasn't just about how hard they pushed, but where they pushed.

The paper also admits a funny mistake they made. At first, they tried to test if the robot was using its "bridge" thoughts by nudging just enough to make the window change. They stopped pushing the moment the window flipped, thinking they had done enough. But because the window is so sensitive, they stopped way too early—before the robot's brain had even started to change. They got a result that said "nothing happened," but it was a false alarm because they didn't push hard enough.

The main takeaway is a warning for anyone trying to read a robot's mind: You can't just look at the window and assume you know what the robot is thinking. The window might change its mind long before the robot does, or it might not change at all even when the robot is doing something huge. To really understand the robot, you have to measure how hard you are pushing and see how the robot reacts at every step, not just at one single moment. The window is a useful tool, but it's not a perfect mirror of the computation happening inside.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →