← Latest papers
🤖 AI

Representation Without Control: Testing the Realization Effect in Language Models

This paper demonstrates that while large language models exhibit prompt-sensitive behavioral shifts and possess linearly decodable internal representations of the "realization effect," the inability to causally steer these representations to alter risk-taking decisions reveals that latent readouts do not necessarily indicate genuine reliance on those representations for downstream decision-making.

Original authors: Ciarán Walsh, Emilio Barkett

Published 2026-05-26
📖 4 min read☕ Coffee break read

Original authors: Ciarán Walsh, Emilio Barkett

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have a very smart, very chatty robot that can pretend to be a human. You ask it questions about money, gambling, and risk, hoping it will act like a person making real-life choices. But here's the big question: Is the robot actually "thinking" like a human, or is it just good at guessing what words to say based on how you asked the question?

This paper puts that robot to the test using a famous psychological trick called the "Realization Effect."

The Setup: The "Open vs. Closed" Wallet

In real life, humans behave differently depending on whether a win or loss is "open" (paper) or "closed" (realized).

  • Paper (Open): Imagine you win $50 on a scratch-off ticket but haven't cashed it in yet. You might feel like you're still "in the game" and take bigger risks to double your money.
  • Realized (Closed): Imagine you cashed that $50 in, put it in your pocket, and then lost it. You might feel like that money is gone and become more cautious.

The researchers asked: Does the AI know the difference between an "open" wallet and a "closed" one, and does it actually change its behavior because of it?

They tested the AI in three ways, like checking a suspect at three different levels of a police investigation.

Level 1: The "Acting" Test (Behavior)

First, they just asked the AI to answer questions.

  • The Result: The AI did change its answers based on whether the story said the money was "open" or "closed." It was sensitive to the words.
  • The Catch: But it didn't act like a real human. When humans lose money in an "open" account, they get reckless. When the AI lost money in an "open" account, it didn't get reckless in the same way. It was just mimicking the shape of the answer, not the logic behind it.

Level 2: The "X-Ray" Test (Reading the Brain)

Next, the researchers looked inside the AI's "brain" (its internal math layers) to see if it actually had a concept of "Open vs. Closed" stored there.

  • The Result: They found it! At a specific layer (Layer 18) of the AI, there was a clear, readable signal. If you looked at the math, you could tell exactly which prompts were "open" and which were "closed."
  • The Metaphor: It's like finding a light switch in a house that is clearly labeled "Kitchen Light." You can see the switch, and you know it's connected to the kitchen.

Level 3: The "Remote Control" Test (Causal Control)

This is the most important part. If the AI really "understands" the concept, then we should be able to flip that internal switch and force the AI to change its mind.

  • The Experiment: The researchers used a "remote control" (activation steering) to manually flip that "Open vs. Closed" switch inside the AI's brain while it was answering questions. They tried to force the AI to act as if it had an "open" wallet or a "closed" wallet.
  • The Result: Nothing happened. Even though they could see the switch and flip it, the AI's final decisions about gambling and risk didn't change.
  • The Metaphor: It's like finding a light switch labeled "Kitchen Light," flipping it on and off, and realizing the kitchen light doesn't turn on. The switch exists, but it's not actually connected to the bulb. It's just a decoration.

The Big Conclusion

The paper concludes that just because an AI can say the right words (Level 1) and has a readable "switch" inside its brain (Level 2), it doesn't mean that switch actually controls its actions (Level 3).

  • The "Realization Effect" in humans is a deep, causal mechanism: The way we feel about money changes how we gamble.
  • The "Realization Effect" in this AI was just a surface pattern. The AI noticed the words "open" and "closed" and adjusted its vocabulary, but it didn't actually use that concept to make its final decision.

The Takeaway:
Don't assume an AI is "thinking" like a human just because it gives human-like answers. It might just be a very good actor. To prove it's actually "thinking," you have to show that you can reach inside, flip a switch, and make it do something different. In this case, the researchers tried to flip the switch, and the AI didn't budge.

In short: The AI has the words for the concept, and it has the internal signal for the concept, but it doesn't have the control over its behavior. The signal is there, but it's not doing any real work.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →