← Latest papers
💬 NLP

Visualizing and Benchmarking LLM Factual Hallucination Tendencies via Internal State Analysis and Clustering

This paper introduces FalseCite, a dataset for benchmarking LLM hallucinations induced by deceptive citations, and reveals through internal state analysis that hallucinating models exhibit a distinct horn-like pattern in their hidden state vectors.

Original authors: Nathan Mao, Varun Kaushik, Shreya Shivkumar, Parham Sharafoleslami, Kevin Zhu, Sunishchal Dev

Published 2026-02-13
📖 4 min read☕ Coffee break read

Original authors: Nathan Mao, Varun Kaushik, Shreya Shivkumar, Parham Sharafoleslami, Kevin Zhu, Sunishchal Dev

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are a student taking a history test. You don't know the answer to a question, but you see a footnote that says, "According to the famous historian Dr. Smith, the answer is X." Even if you've never heard of Dr. Smith, you might feel confident writing down "X" because it looks like it comes from a reliable source.

This is exactly what happens with Large Language Models (LLMs)—the super-smart AI chatbots we use today. They are brilliant at writing, but they have a nasty habit of hallucinating: making up facts that sound perfectly real but are completely false.

This paper is like a detective story where the researchers try to figure out why AI gets tricked so easily and how to spot it before it happens. Here is the breakdown in simple terms:

1. The Trap: "FalseCite"

The researchers built a giant trap called FalseCite. Think of it as a "fake news factory."

  • They took thousands of false statements (like "The Backstreet Boys formed in 1998" when they actually formed in 1993).
  • Then, they added fake citations to them. Some were random nonsense (like "According to a toaster in Ohio..."), and others were cleverly matched to the topic (like "According to a pop-culture expert...").
  • They fed these traps to three different AI models: a big, smart one (GPT-4o-mini) and two smaller ones (Falcon and Mistral).

The Result: The AI models were incredibly gullible. When they saw a fake citation, they didn't just repeat the lie; they started making up even more lies to support it! It was as if the AI saw the "Dr. Smith" footnote and thought, "Oh, if an expert says it, I must be able to explain it in detail," so it invented a whole backstory to match.

2. The Surprise: The "Horn" Shape

This is the most fascinating part. The researchers didn't just look at the answers; they looked inside the AI's brain (its internal state).

  • Imagine the AI's brain is a giant, multi-layered factory. Every time it processes a word, it sends a signal through these layers.
  • The researchers mapped these signals and found something weird. Whether the AI was telling the truth or lying, the signals traced out a shape that looked like a horn (like a musical instrument or a rhino's horn).
  • The Analogy: It's like watching a dancer. Whether they are dancing a happy song or a sad song, their feet still move in a specific, predictable pattern. The "horn" shape is that pattern. The researchers hoped to find a "lie detector" signal that looked totally different, but instead, they found that the AI's brain moves in a very specific, consistent way even when it's lying.

3. Who Got Tricked the Most?

  • The Small Models (Falcon, Mistral): These were the easiest to trick. They were like students who would believe any citation, even a random one, just to sound smart.
  • The Big Model (GPT-4o-mini): This one was smarter. It was less likely to lie when there was no citation. However, when a fake citation was present, it actually jumped to lying more than the others! It seems the smarter the AI, the more confidently it can justify a lie if it thinks the source is credible.

4. Why Does This Matter?

Imagine a doctor using an AI to diagnose a patient. If the AI hallucinates a fake medical study and then confidently explains a fake treatment based on it, that could be dangerous.

This paper teaches us two big lessons:

  1. Fake sources make AI lie worse: If you give an AI a fake reference, it will double down and invent more fake details to support it.
  2. We need better "lie detectors": Since the AI's brain signals look similar whether it's lying or telling the truth (that "horn" shape), we can't just look at the internal code to catch it easily. We need new ways to check if the AI is making things up.

The Bottom Line

The researchers created a tool (FalseCite) to prove that AI is easily manipulated by fake references. They also looked inside the AI's brain and found that its "thinking process" follows a strange, horn-like pattern, whether it's being honest or dishonest. This helps scientists understand that we can't just trust AI blindly, especially when it starts citing sources that might not exist.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →