← Latest papers
💬 NLP

Production and Perception in LLMs: A Token Probability Approach

This study demonstrates that large language models exhibit a distinct production-perception asymmetry in their token probability distributions, where prompt framing alone induces significantly different likelihoods for generated text compared to re-evaluated text, a phenomenon that replicates across diverse model architectures and decays as sequence context accumulates.

Original authors: Anna Marklová, Jiří Milička, Martina Vokáčová, Rudolf Rosa

Published 2026-07-14
📖 5 min read🧠 Deep dive

Original authors: Anna Marklová, Jiří Milička, Martina Vokáčová, Rudolf Rosa

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have a super-smart robot poet. You ask it to write a poem about a hedgehog, and it does. But here's the twist: this robot doesn't just "think" and then "speak" like a human. It uses the exact same brain circuitry to read what you wrote and to write its own words back. It's like a one-way street where the car drives both ways using the same engine.

Because of this, scientists wondered: Does this robot actually have a difference between "writing" a poem and "reading" a poem? Or is it just the same process in reverse?

The Big Discovery: The "Role-Play" Effect

The researchers found that yes, the robot does act differently, but not because it has feelings or intentions. It changes its behavior based on the prompt—the specific way you ask it to do something.

Think of the robot like a method actor.

  • Scenario A (Production): You say, "Please write a poem about a hedgehog." The robot puts on a "Writer" hat. It starts generating words, and it assigns a certain level of confidence (probability) to each word it picks.
  • Scenario B (Perception): You take the exact same poem the robot just wrote, and you say, "Here is a poem about a hedgehog. Please rate it." The robot puts on a "Reader/Critic" hat. It doesn't generate new words; instead, it re-evaluates the words it just saw.

The study measured how much the robot's "confidence" in each word changed between these two scenarios.

The Result: When the robot switched from "Writer" mode to "Reader" mode, its confidence in the words shifted significantly. The difference was huge. In fact, the gap between how the robot "wrote" the poem and how it "read" the poem was about 1.8 times larger than the gap between two different ways of asking it to write the poem.

To put it in numbers:

  • When the robot was asked to write in two slightly different ways (e.g., "Write a poem" vs. "I'd like you to write a poem"), the difference in its word choices was tiny (about 0.010).
  • But when it switched from writing to reading, the difference jumped to 0.034.

This suggests that simply changing the "framing" of the request—telling the robot to be a creator versus a critic—is enough to make its internal math work differently, even though it's using the same brain parts for both jobs.

What the Robot is NOT Doing

The paper is very clear about what this is not.

  • It is not because the robot suddenly gained human-like intentions or a soul. The robot doesn't actually "want" to write or "want" to read.
  • It is not because the robot is better at one task than the other.
  • It is not just a glitch caused by the specific words used in the prompt. The researchers tested this by using four different ways to ask for a poem and three different ways to ask for a rating. The effect happened every single time, no matter how they phrased the request.

The "First Impression" Rule

There's another cool detail about when this happens. The difference between "writing" and "reading" is strongest at the very beginning of the poem.

Imagine the prompt is a loud voice shouting instructions at the start of a race.

  • At the start (Token 1–10): The "Reader" voice is very loud, and the robot's confidence shifts a lot.
  • As the poem goes on: The robot starts reading its own words. The story it's telling (the context) starts to drown out the original instruction. The difference between the "Writer" and "Reader" modes gets smaller and smaller as the poem gets longer.

The researchers modeled this decay and found that the initial "shock" of the prompt fades away, but a tiny bit of difference remains even at the end.

How Sure Are They?

The team didn't just guess; they ran the numbers on 1,089 different poems (about 93,000 words total) using five different robot models (including Llama-3.1-8B, EuroLLM-9B, and others).

  • They found the effect in all five models, including both the "raw" base models and the ones trained to be helpful assistants.
  • The difference was so consistent that the ranges of results for "writing" and "reading" didn't overlap at all.

However, the authors are careful to say this suggests a functional difference, not that the robot has a human-like mind. They note that while the difference is statistically real and measurable, we don't yet know if it's "big enough" to mean the robot is truly thinking like a human who distinguishes between speaking and listening.

The Bottom Line

This study shows that even though Large Language Models use a single, symmetrical engine for both reading and writing, how you talk to them changes how they process the words.

If you tell the robot to be a poet, it thinks one way. If you tell it to be a critic, it thinks another way. It's a bit like a chameleon: the robot's internal "colors" (its probability distributions) shift to match the role you give it, proving that the context of the conversation matters just as much as the words themselves.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →