Language Models Struggle to Use Representations Learned In-Context
Despite evidence that large language models can encode new semantic representations from in-context information, this study reveals that both open-weights and state-of-the-art closed-source models struggle to effectively utilize these learned representations to adapt their behavior for downstream tasks.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
The Big Picture: The "Smart Student" Who Can't Take Notes
Imagine a brilliant student (the AI) sitting in a classroom. The teacher writes a complex, made-up rule on the board using strange symbols (like "ink" means "city" and "toy" means "air"). The student looks at the board, studies it, and seems to understand the pattern perfectly.
The researchers of this paper wanted to know: If the student understands the pattern while looking at the board, can they actually use that understanding to solve a new problem later, once the board is cleared?
The short answer is: No, not really. Even though the student seems to "get it" in the moment, they struggle to apply that knowledge when asked to do something new with it.
The Experiment: Three Different Tests
The researchers ran three main tests to see how well AI models could learn and use these new, made-up rules.
1. The "Memory Lane" Test (Next-Token Prediction)
The Setup: Imagine a game where you walk through a secret city made of random words. You walk from "Toy" to "Ink" to "City." The AI sees this path.
The Twist:
- Scenario A: The AI sees the path written in its own response (like it's writing a diary). It can predict the next word easily.
- Scenario B: The teacher (the user) writes the path, and then asks the AI to predict the next word after a pause.
The Result: In Scenario B, the AI gets stuck. It's as if the student studied the map, but when asked to draw the next street on a blank piece of paper, they forgot the map existed. The AI learned the pattern, but it couldn't "pull it out" of its memory to use it.
2. The "Shape-Shifter" Test (Adaptive World Modeling)
The Setup: The AI is shown a random walk through a grid (like a 4x4 checkerboard) made of random words. Then, the teacher gives a few examples of a new rule, like "Move two steps to the right." The AI has to apply this rule to a new word.
The Result: The AI fails miserably.
- Why? The researchers checked the AI's "brain" (its internal math) and found that it did actually learn the shape of the grid. It knew the map.
- The Problem: Even though the map was stored in its brain, the AI couldn't use that map to solve the puzzle. It's like having a perfect GPS in your pocket but being unable to read the directions when you need to turn. The knowledge was "inert"—it was there, but it was useless for the task at hand.
3. The "Super-Thinker" Test (Reasoning Models)
The Setup: The researchers tried the same tests on the newest, most powerful AI models (the "reasoning models"). These are AIs that are trained to "think out loud" by writing long chains of reasoning before giving an answer.
The Result: These super-models did better, but they still weren't perfect.
- They were great at simple, one-dimensional lines (like a straight road).
- They completely collapsed when faced with 2D grids (like a city block).
- Even when they "thought out loud," they struggled to turn the internal map they had learned into a usable tool. They could describe the map if asked, but they couldn't reliably use it to navigate.
The Core Discovery: "Inert" Knowledge
The main finding of the paper is a concept the authors call "Inert Representations."
Think of it like this:
- Encoding: The AI is like a sponge. When you pour water (new information) on it, it soaks it up immediately. It holds the water.
- Deploying: But when you squeeze the sponge to get the water out to water a plant (solve a new task), the water doesn't come out.
The paper proves that just because an AI can "encode" (store) a new pattern in its brain during a conversation, it doesn't mean it can "deploy" (use) that pattern to solve a problem later. The knowledge is trapped.
What About the "Thinking" AIs?
The researchers hoped that the newest models, which "think out loud" (write long reasoning chains), might fix this problem. Maybe if they talk through the problem, they can unlock the trapped knowledge?
The paper says: Partially, yes, but not really.
These advanced models are slightly better at using the information, but they still fail significantly when the task gets complex (like moving in two dimensions). They are still struggling to turn their "in-context learning" into a flexible tool.
The Conclusion
The paper concludes that while AI models are getting better at recognizing patterns in the moment, they are not yet "adaptable agents." They cannot reliably learn a new rule from a conversation and then immediately use that rule to solve a different problem.
To build truly adaptable AI, we need to figure out how to make that "squeezing" process work—how to make the knowledge the AI learns in the moment actually useful for the tasks it faces next.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.