Beyond Fixed Directions: Adaptive Representation Analysis of Reasoning and Memorization in LLMs
This paper demonstrates that while reasoning and memorization tasks in LLMs are decodable via a single representation direction, this direction is not fixed but undergoes substantial geometric reorganization during reinforcement learning, even as the underlying information remains accessible.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
In the rapidly evolving world of artificial intelligence, large language models have become capable of solving complex puzzles and recalling vast amounts of information. Scientists have long wondered how these digital brains actually work. One prevailing idea suggests that specific behaviors, like logical reasoning or simple fact retrieval, are tied to distinct, fixed pathways within the model's internal memory. Imagine the model's mind as a vast, multi-dimensional space where every thought leaves a trace. If this idea holds true, then the difference between a model thinking through a math problem and one simply recalling a fact should be visible as a single, straight line pointing in a specific direction. This concept is so appealing that some researchers have tried to use these fixed lines as guides to train models to think better, assuming the path they found would remain steady and reliable.
However, a new study challenges the assumption that these paths stay put. Researchers took a closer look at a small language model, specifically a version known as Qwen3-0.6B, to see if this fixed-direction idea holds up when the model is put through rigorous training. They gathered a carefully balanced set of 400 examples, half of which required logical reasoning and half of which required factual recall. By stripping away distractions like sentence length or specific vocabulary, they ensured that any difference they found was truly about the type of thinking involved. Their first discovery was reassuring: they confirmed that, at any single moment in time, the difference between reasoning and remembering is indeed so clear that a single line can separate the two groups perfectly. In fact, this simple line worked just as well as a complex, high-dimensional analysis that looked at every possible angle of the data.
The story takes a turn when the researchers trained the model using a powerful technique called group-relative policy optimization, a method designed to boost reasoning skills. They expected that if the fixed-direction idea were true, the line separating reasoning from memory would stay in the same place, even as the model learned. Instead, they found that while the model still knew the difference between the two tasks perfectly well, the actual direction of that knowledge had shifted dramatically. After training, the line that once pointed clearly toward "reasoning" had rotated so much that it was no longer aligned with its original position. On average, the new direction was only about 45 percent similar to the old one, and in the deepest layers of the model, the shift was even more pronounced, with the internal representation drifting by more than half.
This finding reveals a crucial distinction between what a model knows and how it stores that knowledge. The information itself did not vanish; the model could still distinguish between a math problem and a trivia question with perfect accuracy after training. But the geometric coordinates it used to express that distinction had been completely reorganized. It is as if a map of a city remained perfectly accurate in telling you where the library is, but the compass directions on the map had been rotated so that "north" now pointed where "east" used to be. The researchers measured this shift by comparing the model's internal state before and after training, finding that the change was not just a minor adjustment but a substantial reorganization of the model's internal landscape.
The study concludes that while the ability to separate reasoning from memory is robust, the idea that this ability relies on a single, unchanging direction is incorrect. The information persists, but the specific path it takes through the model's mind is fluid and changes with learning. This suggests that for future efforts to guide artificial intelligence using these internal directions, scientists cannot simply find a path once and assume it will remain valid. Instead, they must accept that the map itself is constantly being redrawn, even as the landmarks remain in place. The stability of the knowledge does not guarantee the stability of the path used to access it.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.