← Latest papers
💬 NLP

From Prerequisites to Predictions: Validating a Geometric Hallucination Taxonomy Through Controlled Induction

This study validates a geometric hallucination taxonomy in GPT-2 by demonstrating that coverage-gap (Type~3) failures are the most distinctive and stable failure mode, characterized by robust norm separation in static embeddings, while confirming that center-drift and wrong-well convergence types do not geometrically separate and that token-level analyses suffer from severe pseudoreplication.

Original authors: Matic Korun

Published 2026-03-03
📖 6 min read🧠 Deep dive

Original authors: Matic Korun

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have a very smart but slightly naive robot (GPT-2) that loves to tell stories. Sometimes, this robot gets confused and makes things up. We call these made-up stories "hallucinations."

For a long time, scientists have tried to figure out how the robot gets confused. A new paper by an independent researcher named Matic Korun proposes a specific map for these confusions. He suggests there are three distinct ways the robot can mess up, and he tried to prove this map is real by forcing the robot into specific traps.

Here is the story of the experiment, explained simply.

The Three Types of Robot Confusion

Korun's map says the robot's mistakes fall into three categories:

  1. Type 1: The "Drifting Drunk" (Center-Drift)
    • The Scenario: You give the robot a prompt with almost no instructions, like just saying "The..." or "It is..."
    • The Mistake: The robot has no idea where to go, so it wanders aimlessly, spitting out the most common, boring words it knows (like "is," "the," "a"). It's like a person staring at a blank wall and mumbling random words because they have no direction.
  2. Type 2: The "Wrong Turn" (Wrong-Well Convergence)
    • The Scenario: You give the robot a tricky sentence that could mean two different things, like "The bank announced..." (Is it a river bank or a money bank?).
    • The Mistake: The robot picks one meaning confidently and sticks to it, even if it's the wrong one for your context. It's like a GPS confidently driving you to the wrong city because the road signs were ambiguous.
  3. Type 3: The "Blank Page" (Coverage Gap)
    • The Scenario: You ask the robot to combine two things it has never seen together, like "The alien biology of a medieval castle."
    • The Mistake: The robot has no data for this. It's not just confused; it's in a territory where its map doesn't exist. It tries to guess, but it's essentially hallucinating because it has no foundation to stand on.

The Experiment: The "20-Run" Test

To see if this map is real, Korun didn't just ask the robot once. He set up a rigorous test:

  • He created 15 different prompts for each of the three confusion types.
  • He ran the experiment 20 times with different random seeds (like rolling dice 20 times) to make sure the results weren't just luck.
  • He looked at the robot's "brain" in two different ways:
    1. Static View: Looking at the dictionary definition of the words the robot chose (ignoring the context).
    2. Contextual View: Looking at the robot's internal "thought process" (the hidden states) right before it chose a word.

The Big Discoveries

1. The "Coverage Gap" (Type 3) is the Easiest to Spot

The map worked perfectly for Type 3. When the robot was asked about impossible combinations (like "alien biology"), its internal "confidence meter" dropped significantly.

  • Analogy: Imagine the robot is a flashlight. When it's talking about normal things, the beam is bright and steady. When it hits a "Coverage Gap," the beam flickers and gets dimmer.
  • The Result: In the "Static" view, the robot's vocabulary choices were clearly different. In the "Contextual" view, the robot's internal confidence (the size of its internal signal) was consistently lower. This was the only type of mistake that showed up clearly on the map.

2. The "Drunk" and the "Wrong Turn" Look the Same

This was the surprising part. The map predicted that Type 1 (Drifting) and Type 2 (Wrong Turn) would look different.

  • The Result: They didn't. In the robot's brain, these two types of confusion looked almost identical. The robot's internal signals for "I'm bored and guessing" and "I'm confident but wrong" were indistinguishable.
  • Analogy: It's like trying to tell the difference between a person who is lost in a fog and a person who is confidently walking in the wrong direction. To the robot's internal sensors, they both just look like "normal walking." The researcher suspects the robot is just too small (124 million parameters) to tell the difference, or the difference is hidden in a part of the brain we aren't looking at yet.

3. The "Fake Significance" Trap (Pseudoreplication)

This is a crucial warning for other scientists.

  • The Problem: When you look at every single word the robot says (thousands of words), the math says the differences are huge and obvious.
  • The Reality: Those words aren't independent. If the robot says "The," it's very likely to say "cat" next. They are linked.
  • The Analogy: Imagine you ask a friend 100 questions, but they just repeat the same answer 100 times. If you count that as 100 different opinions, you are tricking yourself.
  • The Finding: The paper found that looking at individual words makes the results look 4 to 16 times stronger than they actually are. When you look at the prompts as a whole (the real unit of truth), the "obvious" differences for Type 1 and Type 2 vanish.

The "Micro-Signal" Mystery

The paper concludes that the robot's brain is very "saturated." Even when it is making a huge mistake (Type 3), its internal signals don't crash; they just get slightly quieter.

  • Analogy: It's not like the robot's brain is on fire when it hallucinates. It's more like the room gets slightly dimmer. The change is real, but it's so subtle that you need very sensitive equipment (and a lot of data) to notice it.

Why This Matters

This paper is a "reality check."

  1. It validates the map for the hardest errors: We can now detect when a robot is making things up because it has no knowledge (Type 3).
  2. It warns us about the easy errors: We cannot yet easily tell if a robot is just drifting or confidently wrong (Type 1 vs. Type 2) using current methods.
  3. It fixes the math: It tells researchers, "Stop counting every word as a separate vote. You are inflating your results."

In short, the researcher built a better map for robot hallucinations, found that some parts of the map are crystal clear, while other parts are still foggy, and warned everyone not to get fooled by the fog.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →