Verbal Confidence Saturation in 3-9B Open-Weight Instruction-Tuned LLMs: A Pre-Registered Psychometric Validity Screen
This study demonstrates that instruction-tuned open-weight LLMs in the 3-9B parameter range fail psychometric validity tests for verbalized confidence, showing extreme ceiling effects and a lack of correlation between verbalized certainty and actual accuracy or token-level logprobabilities.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
The "Overconfident Student" Problem: Why Your AI is Lying to You About How Sure It Is
Imagine you are a teacher grading a class of students. You ask them a series of difficult trivia questions. After each answer, you ask them: "On a scale of 1 to 100, how sure are you that you're right?"
You expect some students to say, "I'm 90% sure," and others to say, "I'm only 20% sure, I'm just guessing." This allows you to see who actually knows the material and who is just throwing darts in the dark.
This paper is about a group of "students" (small AI models) who have a very strange habit: No matter how hard the question is, they all shout, "I AM 100% SURE!"
The Core Discovery: The "Broken Thermometer"
The researcher, Jon-Paul Cacioli, tested seven different small AI models (the "students"). He wanted to see if their "confidence scores" actually meant anything. In science, if a tool can't tell the difference between a "hot" and "cold" situation, we call it invalid.
The results were startling:
- The Ceiling Effect (The Broken Thermometer): Almost every model the researcher tested gave a confidence score of 95% or higher, regardless of whether they got the answer right or wrong. It’s like having a thermometer that only ever reads "100 degrees." It doesn't matter if you're in a freezer or a sauna; the thermometer just says "Hot!" This makes the score useless for telling if the AI is actually correct.
- The "Multiple Choice" Disaster: The researcher tried to fix this by giving the AI a different way to express confidence (like a multiple-choice scale: No chance, Maybe, Almost certain). Instead of helping, this actually broke the AI's brain. The models became so confused by the new way of answering that they stopped answering the trivia questions correctly altogether.
- The "Thinking Too Much" Paradox: In one specific model that "thinks" before it speaks (a reasoning model), the researcher found that the longer the AI spent "thinking," the less confident it sounded. It’s like a student who mumbles and hesitates for five minutes before finally saying, "I'm 100% sure!"
Why Does This Matter? (The "Safety" Problem)
You might think, "So what? If the AI is always confident, just look at its answer."
But in the real world, we want to use AI for important things—like medical advice, legal research, or driving cars. In those cases, we need the AI to say: "I don't know, I'm not sure, please ask a human."
If an AI is programmed to always sound 100% confident, it will never "abstain" or ask for help. It will confidently drive a car into a ditch or confidently give you the wrong dosage of medicine.
The Takeaway
The paper concludes that small AI models are currently "verbal liars." They might actually know they are wrong deep down in their digital "brains," but when you ask them how they feel, they can't communicate that uncertainty to you. They are stuck in a loop of extreme overconfidence.
The Lesson: Before we trust an AI to make big decisions based on its "confidence," we need to make sure it actually has a working "uncertainty sensor"—and right now, for these models, that sensor is broken.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.