When Context Misleads: Surprisal, Energy and Attention Entropy as Metrics of Coherence Illusions in LLMs
This paper demonstrates that Dutch language models exhibit human-like coherence illusions where distractor contexts reduce surprisal for incoherent continuations, while attention entropy and energy metrics reveal the underlying shared mechanisms driving these effects.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
The Big Idea: When Your Brain (and AI) Gets Tricked
Imagine you are reading a story. Suddenly, you come across a word like "again." Your brain immediately jumps back to the beginning of the story to find out what is happening "again."
Usually, this works perfectly. But sometimes, the story is a trick. There is a "distractor"—a similar-sounding event mentioned earlier that doesn't actually fit, but looks close enough to fool your brain.
The Paper's Discovery:
The researchers found that Large Language Models (LLMs)—the AI brains behind tools like chatbots—fall for this same trick. Just like humans, when an AI sees a confusing story with a matching distractor, it gets confused and thinks the story makes sense when it actually doesn't. They call this a "coherence illusion."
The Experiment: The "Fruit Bowl" Test
To test this, the researchers used Dutch stories (since they had data on how humans read Dutch). Here is a simplified version of the setup:
- The Setup: A character, let's say Megan, is hungry.
- The Twist: The story mentions two things:
- Target: Megan ate a yellow apple.
- Distractor: She did not eat a green pear.
- The Trap: The next sentence says: "The next morning, Megan again ate a pear."
The Human Reaction:
A human reader sees "again" and "pear." Their brain gets confused because the story said she didn't eat a pear. However, because "pear" was mentioned earlier (even though it was negated), the brain sometimes slips up and thinks, "Oh, she must have eaten a pear again!" This is the illusion.
The AI Reaction:
The researchers asked various AI models to read these stories. They found that the AIs behaved exactly like the humans:
- When the story was perfectly logical, the AI was calm.
- When the story was illogical (the trick), the AI got "surprised" (a technical term for "this is unexpected").
- Crucially: When the trick included a matching distractor (mentioning the pear earlier), the AI became less surprised. It fell for the illusion, thinking the confusing story was actually okay.
How They Measured the "Trick"
The researchers didn't just ask the AI if it was confused; they looked inside the AI's "brain" to see how it was thinking. They used three special tools:
1. Surprisal (The "Gasp" Meter)
Think of Surprisal as a "gasp meter."
- If you read a sentence that makes perfect sense, you don't gasp. The meter is low.
- If you read something weird, you gasp. The meter goes high.
- The Finding: When the AI read the trick story, it gasped (high surprisal). But when the trick story had a "distractor" that matched the ending, the AI stopped gasping as much. It was tricked into thinking the story was smooth.
2. Attention Entropy (The "Focus" Flashlight)
Imagine the AI has a flashlight that shines on the words it is thinking about.
- Low Entropy (Focused): The flashlight is a tight beam. The AI is looking at one specific word.
- High Entropy (Diffuse): The flashlight is a wide, fuzzy beam. The AI is looking at many words at once, unsure of which one matters.
- The Finding: When the AI was tricked, the flashlight became fuzzy (high entropy). It was looking at the wrong words because the distractor was confusing it. By turning off the specific parts of the AI responsible for this fuzzy focus, the researchers could see that these parts are the ones getting tricked.
3. Energy (The "Magnet" Test)
This is a new tool the researchers introduced, borrowed from physics. Imagine the AI's memory is a landscape of hills and valleys.
- Low Energy: The AI finds a perfect "valley" where the current word fits perfectly with the past. It's like a magnet snapping into place.
- High Energy: The AI is stuck on a "hill." The word doesn't fit well, and the AI has to struggle to make sense of it.
- The Finding: When the AI was tricked by the distractor, the "energy" dropped. It felt like the word fit perfectly, even though it didn't. This confirmed that the AI was experiencing the same illusion as humans: it felt a false sense of connection.
Why This Matters (According to the Paper)
The paper doesn't say this will fix AI or cure diseases. Instead, it makes a specific, important point:
AI and humans share a similar "glitch."
Both humans and these AI models rely on memory retrieval. When we try to remember what happened earlier in a story, we use clues. If a clue looks similar to the answer (even if it's wrong), our brains (and the AI's math) get hijacked.
The researchers showed that:
- AIs get fooled: They aren't perfect logic machines; they can be tricked by context just like people.
- We can see the trick: By measuring "surprisal," "attention," and "energy," we can see exactly when and how the AI is falling for the illusion.
- It's a shared mechanism: The part of the AI that gets confused is the same part that helps it understand normal stories. It's not a bug; it's a feature of how these systems (and our brains) work.
The Bottom Line
This paper is like a magic trick reveal. It shows that when you give an AI a story with a clever distraction, it doesn't just calculate the answer; it gets "distracted" in a way that mimics human confusion. By studying this, we learn that AI and human brains might be using similar shortcuts to process language, shortcuts that can sometimes lead us both astray.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.