Distilling Self-Consistency into Verbal Confidence: A Pre-Registered Negative Result and Post-Hoc Rescue on Gemma 3 4B
This paper reports that while a pre-registered attempt to distill self-consistency into verbal confidence on Gemma 3 4B failed due to label-entropy collapse, a subsequent post-hoc rescue training on unfiltered data successfully compressed high-accuracy self-consistency signals into a single-pass readout with an AUROC2 of 0.774, demonstrating that preserving label entropy is critical for effective confidence calibration.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
The Big Problem: The Overconfident Robot
Imagine you have a smart robot assistant (an AI model) that answers trivia questions. You ask it, "What is the capital of Peru?" and it says, "Lima." Then you ask, "How sure are you?"
The problem this paper tackles is that this robot is terrible at judging its own confidence. Even when it is completely guessing, it says, "I am 99% sure!" It's like a student who guesses on a test but writes "100% sure" next to every answer, whether they got it right or wrong.
Researchers call this a "degenerate verbal channel." The robot actually knows when it's guessing (it has internal signals), but it refuses to admit it out loud. It just outputs a confident "99%" no matter what.
The Experiment: Teaching the Robot to Be Honest
The researcher wanted to fix this. They tried to teach the robot to say "I'm not sure" when it was actually unsure.
The Setup:
- The Teacher: They used a "self-consistency" method. This means they asked the robot the same question 10 times.
- If the robot got the answer right 10/10 times, the teacher said, "This is an Easy question. You should say you are 95% sure."
- If the robot got it wrong 10/10 times, the teacher said, "This is a Hard question. You should say you are 5% sure."
- If it was mixed (e.g., 5 right, 5 wrong), the teacher said, "You are 50% sure."
- The Goal: Train the robot to look at a question and immediately say the right confidence level (5% or 95%) without needing to ask itself 10 times.
The First Attempt: The "Perfect Student" Trap (The Negative Result)
The researcher started with a strict rule (a "modal filter"). They decided to only show the robot examples where the robot was mostly right. They thought, "Why teach the robot about being wrong? Let's only teach it when it's doing well."
What happened?
The robot failed completely. It kept saying "95% sure" for everything.
Why?
Think of it like a music teacher who only lets a student practice songs they already know perfectly. The student never learns how to handle a difficult song.
In the training data, almost every question was an "Easy" one (where the robot was 100% right). The robot learned a simple trick: "If I see a question, just say 95%." It never learned what "low confidence" felt like because it was never shown examples of being unsure. The researchers called this a "Label-Entropy Collapse." The data had no variety, so the robot had nothing to learn.
The Rescue: Letting the Robot See the Mess (The Positive Result)
The researcher realized their mistake. They went back and said, "Okay, let's show the robot everything, including the questions it got wrong."
They removed the strict filter and trained the robot on all 2,000 questions, including the ones where the robot was confused (getting 0 or 1 right out of 10 tries).
The Result:
Suddenly, the robot learned!
- Before: It said "95% sure" for everything.
- After: It started saying "5% sure" when it was guessing and "95% sure" when it knew the answer.
- The Magic: It could now tell the difference between a right and wrong answer just by listening to its own confidence level. It was almost as good as if it had asked itself the question 10 times, but it did it in just one go.
The Side Effect: The Robot Got Smarter (and Quieter)
There was a surprising bonus. When the robot learned to be honest about its confidence, it also got better at answering questions on a different test (MMLU).
- Why? The researchers noticed a change in how the robot spoke.
- Before: The robot would ramble on with long, confusing explanations.
- After: The robot started giving short, direct answers (e.g., "c. The answer is X. Confidence: 95%").
- The Analogy: It's like a student who, when told to be honest about their knowledge, stops trying to "fake it" with a long, rambling essay and just gives the straight answer. The honesty forced the robot to be more concise, which accidentally helped it get more questions right.
The Catch (Limitations)
The paper is very honest about what it didn't do:
- It's Binary: The robot learned to say "5%" or "95%." It didn't learn to say "72%." It's like a light switch (On/Off) rather than a dimmer switch.
- It's a Rescue Mission: The success happened after the researcher realized the first plan was flawed. It needs to be tested again to be sure it works every time.
- Accuracy Drop: On the specific test used for training, the robot actually got slightly worse at answering correctly, even though it got better at judging its confidence. It traded some accuracy for honesty.
The Two Big Lessons
The paper concludes with two main takeaways for anyone trying to teach AI to be honest:
- You must show the robot when it's wrong. If you only train it on "easy" examples where it's confident, it will never learn to say "I don't know."
- Honesty changes the style. Teaching a robot to report its confidence correctly seems to make it stop rambling and start giving cleaner, more useful answers.
In short: The researcher tried to teach a robot to admit when it was guessing. The first attempt failed because they hid the hard questions. The second attempt succeeded, teaching the robot to be a "binary" truth-teller that is much better at spotting its own mistakes than before.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.