← Latest papers
🤖 machine learning

Reported Confidence in LLMs Tracks Commitment More Than Correctness

This paper demonstrates that verbal confidence reports in large language models primarily reflect an internal "commit-readiness" state driving the decision to answer rather than the actual correctness of the response, distinguishing them fundamentally from log-probabilities which track answer evidence.

Original authors: Dharshan Kumaran

Published 2026-06-30
📖 5 min read🧠 Deep dive

Original authors: Dharshan Kumaran

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are asking a very smart, but slightly mysterious, robot assistant to solve a riddle. You ask it, "What is the capital of France?" It thinks for a moment and says, "Paris." Then, you ask it, "How sure are you?"

The robot replies, "I am 100% certain."

In the world of Artificial Intelligence (AI), we usually assume that when the robot says "100% certain," it means, "I have checked my internal database, and I am statistically sure this answer is correct."

However, a new study from Google DeepMind suggests we might be misunderstanding the robot's confidence. The researchers found that when an AI says it is "confident," it isn't necessarily saying, "My answer is right." Instead, it is often saying, "I am ready to stand by this answer, even if it might be wrong."

Here is a breakdown of their findings using simple analogies.

The Two Types of "Confidence"

The study compared two different ways the robot measures its own certainty:

  1. The "Math Confidence" (Log-Probability): This is like a calculator. It looks at the raw numbers and probabilities of its options. If the math says "Paris" is 99% likely to be the answer, the calculator is confident.

    • The Finding: This type of confidence is a good judge of truth. If the math says it's confident, the answer is usually right. If the math says it's unsure, the answer is usually wrong.
  2. The "Verbal Confidence" (The Robot's Words): This is when you ask the robot directly, "How confident are you?" and it answers in words or a number scale (like "Very Sure").

    • The Finding: This type of confidence is a bad judge of truth, but a fantastic judge of commitment. When the robot says, "I am very sure," it doesn't mean the answer is definitely correct. It means the robot has made up its mind and is willing to commit to that answer in front of a user.

The "Commit or Quit" Game

To figure this out, the researchers set up a two-stage game, inspired by how scientists study decision-making in animals:

  • Stage 1: The robot answers a question and gives a confidence rating (e.g., "I'm 90% sure").
  • Stage 2: The robot is shown its own answer and asked: "Do you want to Commit (show this answer to a human) or Abstain (say 'I don't know')?"

The Surprise Result:
The robot's verbal confidence in Stage 1 was a much better predictor of what it would do in Stage 2 than it was a predictor of whether the answer was actually right.

  • If the robot said, "I'm super confident," it almost always chose to Commit in Stage 2, even if the answer was actually wrong.
  • If the robot said, "I'm not sure," it almost always chose to Abstain in Stage 2.

The Analogy:
Imagine a gambler at a casino.

  • Math Confidence is like the gambler checking the odds on a piece of paper. If the odds are good, the gambler knows they are likely to win.
  • Verbal Confidence is like the gambler shouting, "I'm going all in!"
    The study found that the gambler shouting "I'm going all in!" is often just a signal that they are ready to bet, not a guarantee that they have the winning hand. They might be bluffing or just feeling bold. The "Math Confidence" (the odds) is the only thing that actually tells you if they will win.

The "Commit-Readiness" Switch

The researchers dug deeper into the robot's "brain" (its internal code) to see what was happening. They found a specific moment right after the robot gives an answer, but before it is asked to decide whether to commit.

At this exact moment, the robot's internal state is organized around one question: "Am I ready to stand by this?"

  • The robot has a "Commit/Abstain" switch that is very loud and clear in its brain.
  • The robot has a "Right/Wrong" switch that is much quieter and fuzzier.

When the robot generates a verbal confidence report, it is essentially reading the loud "Commit" switch, not the quiet "Right/Wrong" switch. It's like a person who feels a strong urge to speak up; they might feel very confident speaking, even if what they are saying is factually incorrect.

Why This Matters

For a long time, people have treated an AI's verbal confidence ("I'm 95% sure") as a direct measure of accuracy. We assumed that if the AI says it's sure, we can trust the answer.

This paper argues that this is a mistake.

  • Verbal Confidence is a behavioral signal. It tells you: "I am willing to show this answer to a human."
  • Math Confidence is a truth signal. It tells you: "This answer is likely correct."

If you rely on the robot's words to decide if an answer is true, you might get fooled. The robot might be very eager to commit to a wrong answer because its internal "commitment switch" is flipped, even though the "truth switch" is flickering.

The Bottom Line

The study concludes that we need to stop treating an AI's verbal confidence as a simple "truth meter." Instead, we should understand it as a "commitment meter." It tells us when the AI is ready to take a stand, but it doesn't necessarily tell us if that stand is on solid ground. To know if the answer is actually right, we need to look at the underlying math (the log-probabilities), not just the robot's enthusiastic words.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →