← Latest papers
💬 NLP

Does Bielik Know What It Doesn't Know? Activation Dispersion Separates Entity Familiarity from Factual Reliability Across Model Scale

This paper demonstrates that while large language models possess an internal, activation-based signal that accurately distinguishes between known and fabricated entities across various scales, this "awareness" does not translate into behavioral reliability, as models frequently hallucinate about known entities and almost never abstain from answering, revealing a fundamental disconnect between entity familiarity and factual correctness.

Original authors: Grzegorz Brzezinka

Published 2026-07-09
📖 5 min read🧠 Deep dive

Original authors: Grzegorz Brzezinka

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

The Big Question

Imagine you are asking a very smart, well-read robot (an AI) a question. You want to know: Does the robot "feel" different inside when it is talking about something it actually knows, versus when it is making something up?

The researchers wanted to see if they could peek inside the robot's brain (its internal computer signals) before it even finishes its sentence to tell if it is confident or if it is about to hallucinate (lie).

The Experiment: The "Polish Robot"

The team tested a family of Polish AI models called Bielik. They made the robots answer questions about four types of things:

  1. Athletes (Famous ones like Lewandowski, obscure ones, and fake ones).
  2. Cities (Warsaw, small villages, fake names).
  3. Writers and Musicians.

They created three types of questions for each:

  • Known: Famous things the robot definitely knows.
  • Obscure: Real things the robot has probably never heard of.
  • Fake: Made-up names that sound real (like "Roman Lewandowicz").

The Discovery: The "Brain Glow"

The researchers looked at the robot's internal electrical activity (called activations) right after it read the question but before it started typing the answer. They used two simple math tools to measure how "spread out" or "focused" the brain activity was.

The Analogy: The Library vs. The Confused Crowd

  • When the robot knows the answer: Imagine a librarian walking straight to one specific shelf, picking up one specific book, and bringing it to the counter. The activity is focused and concentrated.
  • When the robot doesn't know (or is making it up): Imagine a confused crowd running around the library, touching random shelves, and shouting in all directions. The activity is scattered and messy.

The Result:
The math tools could tell the difference between the "focused librarian" and the "confused crowd" with 95% to 100% accuracy.

  • This worked for robots of all sizes (from small 1.5 billion parameter models to large 11 billion ones).
  • It worked for all types of topics (athletes, cities, etc.).
  • It worked instantly, with just one look at the brain, no extra testing needed.

The Catch:
The robot knows it doesn't know the answer inside its brain, but it doesn't say so. It just keeps talking anyway.

The Two Different "Growth Curves"

This is the most surprising part of the paper. The researchers found two things happening at different speeds as the robots got bigger:

  1. Knowing "I Don't Know" (The Brain Signal):

    • Even the smallest robot (1.5B) had a perfect "I don't know" signal in its brain. It knew immediately when it was being asked about a fake person.
    • Making the robot bigger didn't really change this; it was already at the top level.
  2. Actually Giving the Right Answer (Behavior):

    • The smallest robot was terrible at giving correct facts. It would guess and get it wrong.
    • As the robots got bigger, they got much better at giving correct facts. The biggest robot (11B) was much more accurate.

The Analogy:
Imagine a student taking a test.

  • Small Student: Knows they are guessing (their brain is nervous), but they write down a wrong answer anyway.
  • Big Student: Knows they are guessing (brain is nervous), and they have learned enough to actually write the right answer.
  • The Gap: The small student's brain knew they were guessing, but their mouth kept talking nonsense. The big student's brain knew, and their mouth finally got the facts right.

The "Refusal" Problem

The researchers checked if the robots would ever say, "I don't know" and stop answering.

  • Out of 2,520 answers, the robots only refused to answer twice.
  • Even though their brains were screaming "This is fake!" (with 99% certainty), their mouths kept making up stories.
  • Only the biggest robot ever refused, and even then, it was very rare.

What This Means (In Simple Terms)

  1. We can detect lies: We can build a simple "lie detector" for AI that checks its brain activity before it speaks. If the brain activity looks "scattered," the AI is likely about to make something up.
  2. It's not about size: The ability to know you are guessing is there even in small models. The ability to actually know the facts takes a lot more size and training.
  3. The AI is stubborn: The AI knows it's guessing, but it doesn't stop. It's like a person who knows they are bluffing in a poker game but keeps playing anyway because they are programmed to always give an answer.

Summary

The paper proves that AI models have a clear internal signal that says, "I have never seen this before." This signal is perfect and works instantly. However, the models are very bad at listening to that signal and stopping themselves from making things up. They know they don't know, but they keep talking anyway.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →