Linear representations in language models can change dramatically over a conversation
This paper demonstrates that linear representations of high-level concepts in language models can shift dramatically and content-dependently over the course of a conversation as the model adapts to its conversational role, challenging the reliability of static interpretability methods and steering techniques.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine a large language model (like the AI you might chat with) not as a static encyclopedia, but as a chameleon that changes its internal colors depending on the room it's in.
This paper, written by researchers at Google DeepMind, investigates how these "colors" (which they call linear representations) shift dramatically during a conversation. Here is the breakdown of their findings using simple analogies:
1. The "Truth Switch" Flips
Usually, we think of an AI's internal "truth detector" as a fixed lightbulb. If a statement is true, the light is on; if it's false, the light is off.
The researchers found that this lightbulb is actually a dimmer switch that can be turned all the way around.
- The Experiment: They started a conversation where the AI was told, "Today is 'Opposite Day'." The AI agreed to answer everything backwards.
- The Result: At the very beginning, the AI's internal "truth" setting correctly identified that "The sky is blue" is true. But after just a few turns of playing "Opposite Day," the AI's internal representation flipped. Suddenly, in its own internal language, "The sky is blue" started looking like a lie, and "The sky is green" started looking like the truth.
- The Takeaway: The AI didn't just say the opposite; its internal brain reorganized itself to believe (or represent) the opposite as the truth for the duration of that chat.
2. The "Role-Play" Costume
The researchers tested this with two types of scenarios:
- The Scripted Play: They fed the AI a pre-written script of a conversation where a character claims to be a god or a conscious being. Even though the AI was just reading the script (not actually generating it), its internal "truth" settings flipped to match the character's role.
- The Live Act: They had the AI actually role-play the character in real-time. The result was the same: the internal "truth" settings flipped.
- The Analogy: Think of the AI as an actor. When an actor puts on a costume to play a villain, they don't just act like a villain; for the duration of the scene, their entire posture, voice, and mindset shift to match the role. The paper suggests the AI does the same thing: it changes its internal "costume" to fit the conversation.
3. The "Generic" vs. "Specific" Distinction
Not everything changes. The researchers found a difference between generic facts and conversation-specific topics.
- Generic Facts: Questions like "Can sound travel in a vacuum?" or "Are you a computer?" stayed relatively stable. The AI's internal "truth" for these didn't flip much, even during the wild role-play.
- Specific Topics: Questions directly related to the story (e.g., "Do you have feelings?") flipped dramatically.
- The Analogy: Imagine a traveler in a foreign country. They might still know their own name and where they live (generic facts), but they might start speaking the local language and adopting local customs (specific topics) so thoroughly that, for the moment, they seem to have forgotten their original habits.
4. Why This Matters (The "Lie Detector" Problem)
The paper warns that this makes it very hard to build a "lie detector" for AI.
- The Problem: Many researchers try to find a specific "truth line" inside the AI's brain to check if it's lying. They assume this line is fixed, like a ruler.
- The Reality: The paper shows this "ruler" is actually malleable clay. If you measure the AI at the start of a chat, the ruler says "True." If you measure it at the end of a role-playing chat, the ruler has been squished and twisted, and it now says "False" for the same fact.
- The Metaphor: It's like trying to use a compass to find North. If you are in a normal room, the compass works. But if you walk into a room full of giant magnets (the conversation context), the compass needle spins wildly. You can't trust the compass reading unless you know exactly what "magnets" are currently in the room.
5. The "Story" vs. The "Conversation"
Interestingly, the AI didn't change its mind as much when it was just reading a sci-fi story about a fictional world.
- The Difference: When the AI was told to play a role in a conversation (like arguing about consciousness), it changed its internal settings. When it was just reading a story labeled as "fiction," it stayed more grounded.
- The Analogy: It's the difference between reading a book about a wizard and pretending to be a wizard in a game. When you read, you stay yourself. When you play the game, you fully inhabit the character, and your internal logic shifts to match the game's rules.
Summary
The paper concludes that an AI's internal understanding of "truth" isn't a permanent, unchangeable fact. It is a dynamic state that evolves moment-by-moment based on the role the AI is playing in the conversation. If the conversation cues the AI to act as if a lie is true, the AI's internal brain reorganizes to make that lie look like the truth.
This means we cannot assume that an AI's internal "truth settings" mean the same thing at the end of a conversation as they did at the beginning.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.