Behaviour Is an Incomplete Measure of Reasoning Development: Cross-surface pre-arrival accessibility and the limits of developmental inference in a recurrent-depth reasoner
This paper demonstrates that in a recurrent-depth reasoner, internal hidden-state probes can detect reasoning capabilities long before they manifest in observable behavior, revealing that behavioral thresholds and decoder accessibility are distinct, incomplete measures of the actual developmental timeline of reasoning.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
In the rapidly evolving field of artificial intelligence, researchers often face a fundamental question: when does a machine actually learn something? For years, the standard answer has been behavioral. If a computer program can correctly answer a question it has never seen before, we assume it has acquired the reasoning skill required to do so. We look at the output, the final answer, and declare the capability present. However, this approach treats the mind of the machine as a black box, assuming that the moment the correct answer appears, the internal process of figuring it out must have just happened. This paper challenges that assumption by looking inside the machine while it is still learning, asking whether the ability to solve a problem exists internally long before the machine is capable of showing it on the outside. The researchers are not just watching what the machine does; they are measuring what the machine knows, even when it cannot yet prove it.
To understand the stakes, imagine trying to learn a new language. You might spend months memorizing vocabulary and grammar rules, yet still fail to hold a conversation. A teacher watching only your speech might conclude you have learned nothing. But if you could peek at your brain, you might find that the neural pathways for the language are already forming, waiting for a final spark to connect them. This is the gap the researchers investigated. They built a specialized computer model designed to solve logical puzzles involving chains of relationships, such as figuring out who is related to whom in a family tree. They trained this model using two different methods: one where the relationships were described using abstract, meaningless symbols, and another where the same relationships were described using English-like sentences. The goal was to see if the way the information was presented changed how quickly the model learned, and whether the model's internal state revealed knowledge that its behavior did not.
The results were striking and counterintuitive. When the model was trained on the abstract symbols, it mastered the three-step logical puzzles in just seventy training cycles. When the exact same model, with the exact same training plan, was trained on the English-like sentences, it took over thirteen thousand cycles to reach that same level of competence. That is a difference of more than one hundred and eighty times. Yet, once the sentence-trained model finally cracked the three-step puzzle, it solved the four-step version in just eight cycles. This suggests that the model spent thousands of cycles grinding through a difficult phase, seemingly making no progress, before suddenly unlocking the ability to solve harder problems. If a researcher had only watched the final answers, they would have seen a machine that failed for a long time and then suddenly succeeded, missing the entire hidden journey of internal development that happened during the struggle.
To see what was happening during those long, difficult cycles, the researchers looked inside the model's hidden states—the internal representations of the data that the machine uses to think. They used a simple tool, a linear probe, which acts like a decoder, to check if the correct answer was already present in the machine's mind before the machine actually outputted it. On the sentence-trained surface, this probe could detect the identity of the future answer with a small but reliable signal long before the model could answer correctly on its own. The probe found that information about the answer was accessible internally, even though the model could not yet express it. This internal accessibility was not a fluke; when the researchers switched back to the abstract symbols, they found the same pattern of early accessibility at specific points in the machine's processing. The machine was holding information about the answer internally, waiting for the right moment to release it.
However, the study also uncovered a significant limitation in how we try to measure this kind of learning. The researchers wanted to track exactly how this internal knowledge grew stronger over time, creating a map of the machine's learning curve. They quickly realized, however, that this specific measurement was impossible to perform accurately. The problem was that the group of questions the machine was being tested on kept changing. As the machine learned, it got better at answering certain questions, and those questions were removed from the test group. This meant that the very act of measuring the machine's progress was changing the population of questions being measured. It was like trying to measure how fast a runner is improving by only timing them on the tracks they haven't finished yet; as they get faster, the tracks change, making it impossible to compare their speed fairly over time. The researchers proved that this specific type of measurement is fundamentally flawed in this context, not because the machine didn't learn, but because the method of tracking it was broken.
The paper concludes that we cannot rely on a single way of knowing what a machine has learned. Behavioral success, the ability to give the right answer, is just one observable fact. Internal accessibility, the ability to detect the answer inside the machine, is another. And the actual process of development over time is a third, which often cannot be measured directly if the measurement method itself alters the conditions. The researchers found that a machine can possess a capability internally long before it displays it, and that the path to that capability can look completely different depending on how the problem is presented. Most importantly, they showed that simply observing the machine, no matter how closely, is not enough to understand the computation it has acquired. To truly know what the machine has learned, we must move beyond observation and begin to intervene, testing the machine's internal states directly to see if they are the actual cause of its behavior. Until then, the internal life of these systems remains partially hidden, even when they are finally getting the right answers.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.