Graded Entity-Familiarity Readouts in Language Models: Polish Adaptation, Cross-Language Robustness, and Refusal Steering
This paper demonstrates that language models, particularly those adapted for Polish, possess graded internal readouts of entity familiarity that are robust to prompt language and can be effectively steered to control refusal behaviors, though these pre-generation signals remain distinct from the policies governing final abstention decisions.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
The Inner Monologue of a Robot: When Does It Know It Doesn't Know?
Imagine you are asking a very smart robot a question. Sometimes, the robot knows the answer and speaks up confidently. Other times, it might make up a story because it doesn't actually know the truth, a mistake we call "hallucinating." Scientists have long wondered: Can the robot tell the difference between "I know this" and "I'm guessing" before it even starts typing?
To understand this, we need to look inside the robot's brain, which is made of layers of mathematical connections called "activations." Think of these activations like the electrical signals in a human brain when you see a face. If you see a friend, the signal is strong and specific; if you see a stranger, it's weaker or fuzzy. Researchers have found that these electrical signals often contain a hidden "confidence meter." If the robot is asked about something it has seen many times in its training data, the signal looks different than if it's asked about something totally made up. The big question is: Can we read this meter to stop the robot from lying before it says a single word?
The Polish Detective and the "Familiarity" Meter
This paper is like a detective story, but instead of solving crimes, the author, Grzegorz Brzezinka, is investigating how twelve different AI models "feel" about Polish names and places. The goal was to see if these models have a built-in "familiarity meter" that tells them how well they know an entity, and if that meter works even when the question is asked in a different language.
The Setup: A Library of Polish Names
To test this, the author built a special test set of 1,440 Polish entities. Imagine a library with four sections: athletes, cities, writers, and musicians. In each section, there are real people and places, ranging from super-famous stars (like the top 10% of most-viewed Wikipedia pages) to obscure locals (the bottom 10%). Crucially, the author also invented 240 fake names that looked and sounded exactly like real Polish names but didn't exist. This was the control group: if the robot gets confused by a fake name, it's a liar; if it knows the difference, it's honest.
Finding 1: The Meter is a Dial, Not a Switch
The first big discovery is that familiarity isn't just a light switch (On/Off). It's a dimmer switch. The study found that for models specifically adapted to Polish (like the Bielik and PLLuM families), the "familiarity score" rises smoothly as the entity becomes more popular.
- The Analogy: Imagine a volume knob. For a super-famous singer, the knob is turned all the way up. For a local band, it's halfway. For a made-up name, it's at zero.
- The Surprise: This "dimmer" effect was much stronger in the Polish-adapted models than in the giant, general-purpose models (like Gemma-4 or Qwen3). Even though the general models were bigger (had more "parameters"), they didn't have a better familiarity meter. It turns out that learning Polish specifically was more important for this skill than just making the model bigger. The Polish-adapted models showed a clear correlation (between 0.28 and 0.57) between popularity and their internal confidence, while the others were barely above zero.
Finding 2: The Meter Works Across Languages
The author then asked a tricky question: Is this meter reading the Polish words or the actual entity? To test this, they took a question about a famous Polish athlete and asked it in English ("Who is...?") instead of Polish ("Kim jest...?").
- The Result: The meter barely blinked. The models retained 96% to 101% of their accuracy when switching languages. This suggests the robot isn't just recognizing Polish grammar; it's actually recognizing the person inside the question, regardless of the language used to ask about them.
Finding 3: We Can Hack the "Refusal" Button
This is the most dramatic part of the story. One of the models, Gemma-4-12B, had a habit of saying "I don't know" (refusing) when it was unsure. The author decided to see if they could control this behavior by "steering" the robot's brain.
- The Experiment: They found a specific mathematical direction in the model's brain that represented "familiarity." By adding a tiny nudge in this direction, they could force the robot to change its mind.
- The Magic:
- If they nudged the robot to feel less familiar with a famous person, the refusal rate jumped from 24% to 100%. The robot suddenly decided it didn't know the answer, even though it did.
- If they nudged the robot to feel more familiar with a fake name, the refusal rate dropped from 73% to 0%. The robot confidently started making up biographies for people who didn't exist.
- The Takeaway: This proves that the familiarity signal isn't just a passive observation; it's a causal lever. If you push the "familiarity" button, the robot's decision to speak up or stay silent changes immediately.
Finding 4: The "Pre-Flight" Check
Finally, the paper asked: Can we use this meter to stop the robot from lying before it generates an answer?
- The Comparison: The author compared their "one-pass" meter (which checks the question before the answer is written) against other methods that wait until the answer is finished to check for errors.
- The Verdict: The pre-flight meter was excellent at spotting fake names, beating almost all other methods. However, predicting whether an answer would be factually wrong was harder. The meter worked well for the Polish models, but for the model that refused to answer (Gemma-4), the results were tricky because the model's own habit of saying "I don't know" messed up the error count.
- The Big Picture: The study concludes that while the robot's brain does contain a graded familiarity signal, not all robots use it to decide whether to speak. The Polish models had the signal but never refused to answer. The Gemma model refused, but its signal was less precise. The paper suggests that we might need to build an external "gatekeeper" that reads this signal and decides for the robot whether to answer or look up the info, especially for obscure, long-tail questions.
Why This Matters
This research suggests that AI models have a hidden "gut feeling" about what they know, and we can read it. It's not just about making models bigger; it's about how they are trained. If we can teach models to trust this internal meter, or build tools that listen to it, we might be able to stop them from confidently making things up about things they've never seen. It's a step toward making AI not just smarter, but more honest about its own limits.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.