Are You Sure You're Sure? On the Impact of Instruction Tuning on Confidence and Lexical Diversity
This paper investigates the impact of instruction tuning on language models, finding that while it consistently increases verbalized confidence and reduces cross-rationale diversity, it produces variable effects on surface-level lexical diversity and degrades likelihood-based calibration without significantly altering predictive accuracy.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are talking to a very smart, very well-read robot. You ask it a question, and it answers with a confident voice, explaining exactly why it thinks it's right. But here's the catch: sometimes, that robot is just pretending to be sure. It might be guessing wildly but sounding like a professor. This is a big deal in the world of Artificial Intelligence, specifically for "Large Language Models" (LLMs)—the super-smart computers that write stories, solve math problems, and answer trivia.
To understand this paper, you need to know two things. First, Instruction Tuning. Think of a raw AI model as a brilliant but chaotic student who knows a million facts but doesn't know how to follow a teacher's rules. "Instruction tuning" is like sending that student to a special boot camp where they learn to listen to prompts, follow directions, and answer questions in a specific format. It usually makes them much better at doing what we want. Second, Confidence and Diversity. When an AI answers, it can tell us how sure it is (confidence). It also generates "rationales"—little explanations of its thinking. "Lexical diversity" is just a fancy way of saying "how many different words and sentence structures it uses." If an AI is truly thinking hard, it might try different ways to explain its answer. If it's just reciting a script, it might sound the same every time.
Why do we care? Because in real life—like when a robot doctor diagnoses a patient or a robot lawyer reviews a contract—we need to know if the robot is actually right or just sounding confident. If the robot gets louder and more repetitive without actually getting smarter, that's a dangerous trap.
The "Overconfident Student" Experiment
In this study, the researchers decided to play detective. They wanted to see what happens to these AI students after they finish their "Instruction Tuning" boot camp. Specifically, they asked: Does the robot become more confident because it actually learned more, or is it just acting more confident while its explanations become boring and repetitive?
They took three different families of AI models (think of them as three different schools: Qwen, Mistral, and Llama) and compared their "before" and "after" versions. They asked them thousands of questions from science tests, general knowledge quizzes, and logic puzzles.
Here is what they found, and it's a bit of a plot twist.
1. The Robot Gets Louder, But Not Necessarily Smarter
After instruction tuning, the robots became significantly more confident. They stopped hedging their bets. If you asked them, "How sure are you?" they would say, "I'm 90% sure!" instead of "Maybe 50%."
- The Catch: This confidence didn't always match their actual accuracy. For example, on one test (ARC-Easy), the Llama model's confidence jumped from about 49% to 90%, but its actual score on the test stayed exactly the same at 82.2%. It was like a student raising their hand and shouting "I know this!" with 100% certainty, even though they hadn't actually improved their test score. The researchers found that this "verbalized overconfidence" happened consistently across all the models they tested.
2. The "Scripted" Explanation
The researchers then looked at the explanations (the rationales) the robots wrote. They measured two things:
- Cross-rationale diversity: If you ask the robot the same question five times, does it give you five different explanations, or does it just copy-paste the same one?
- Surface-level diversity: How many unique words and word pairs does it use in a single explanation?
The results were mixed but revealing. The robots became much more repetitive when explaining their answers. If you asked the same question five times, the instruction-tuned models gave very similar explanations, whereas the "raw" models were more varied. It's as if the boot camp taught them to stick to a single, safe script.
- However, the "surface-level" diversity (how many different words they used) was a bit chaotic. Sometimes the robots used more unique words, sometimes fewer. It depended on which robot and which test you were using. There was no single rule for this.
3. The Mismatch
The most important finding is that these two changes—getting more confident and getting more repetitive—didn't always happen together in a neat package.
- Sometimes the robot became less uncertain (more confident) and also less diverse in its explanations.
- Other times, it became less uncertain but more diverse in its word choice.
- The researchers found that you couldn't just look at how many different words a robot used to guess how confident it was. The two things were dancing to different tunes.
4. It's Not Just About Length
You might think, "Maybe the robots just got more confident because they started writing longer explanations?" The researchers checked this. They forced the robots to write explanations of the exact same length and only looked at the ones where they picked the same answer. Even then, the pattern held: the instruction-tuned models were still more confident and still more repetitive in their explanations. The change wasn't just a side effect of writing more or less; it was a fundamental shift in how the models behaved.
The Bottom Line
The paper concludes that instruction tuning makes AI models act like overconfident students who have memorized a single, safe way to explain things. They sound more certain, but they don't necessarily know more, and they stop trying to explain things in different ways.
The researchers suggest that this is a warning sign. If we rely on an AI's confidence to tell us if it's right, we might get fooled. The robot might be shouting "I'm sure!" while actually just repeating the same old script. They also found that the robot's confidence didn't always match its actual accuracy, and its explanations became less varied, which might make it harder to spot when the robot is making a mistake.
In short: Just because the robot sounds more sure of itself after training doesn't mean it's smarter. It might just be better at sounding like it knows what it's doing, even if it's using the same old lines.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.