When LLMs Agree, Are They Right? Auditing Self-Consistency and Cross-Model Agreement as Confidence Signals
This paper challenges the reliability of using self-consistency and cross-model agreement as confidence signals for Large Language Models, demonstrating through a large-scale study that while agreement is a positive but weak predictor of correctness, it often leads to over-confidence in frontier models where high agreement coincides with frequent errors.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are running a massive talent show to find the best AI brain. You have a panel of judges (the AI models) and you ask them a million tricky questions. The big rule everyone follows is: "If all the judges agree on an answer, it must be right." It feels like common sense, right? If ten people all say the sky is green, maybe we're all seeing a weird filter? But if a hundred people say it, it's probably true!
This paper is like a detective story that goes behind the scenes to check if that rule actually works. The author, Kaihua Ding, ran a huge experiment with 53 different "runners" (people running the AI tests) and asked them to generate 265,000 answers to some of the hardest math and science questions available (called GPQA Diamond and AIME).
Here is the twist the paper uncovers: Agreement is not the same thing as being correct.
The "Echo Chamber" Trap
Think of the AI models like a group of students who all studied from the same slightly flawed textbook. If they all memorized the same wrong answer to a question, they will all raise their hands and agree on that wrong answer. They aren't agreeing because they are smart; they are agreeing because they share the same bias or a "memorized heuristic."
The paper found that when AI models agree with themselves (or each other), it's often just them echoing a shared mistake, not a sign of truth.
The "Over-Confident" Star Student
The most surprising discovery involves the "frontier" models—the super-smart, most advanced AI versions. You might think the smartest model would be the most reliable. The paper shows the opposite is true in this specific setup.
- The Finding: The most advanced model (gpt-4.1) agreed with itself on 89% of the time (a very high confidence score). But when it agreed, it was only right about 48% of the time.
- The Analogy: Imagine a student who is so sure of their answer that they shout it out with 100% confidence, but they are actually wrong half the time. The paper calls this "over-confident." In fact, this super-smart model was worse at signaling when it was right than a slightly less advanced "mid-tier" model.
- The Proof: The study measured this using a statistical link (called a correlation) between "agreement" and "correctness." For the smartest model, this link was very weak (around 0.20), meaning its high confidence was almost useless as a trust signal.
Does "Chain-of-Thought" Help?
Some people think if you ask the AI to "show its work" (a method called Chain-of-Thought), it will be more honest.
- The Result: Yes, asking the AI to show its work did make it slightly more accurate (it got more questions right).
- The Catch: However, it didn't make the "agreement signal" much better. The AI still couldn't reliably tell you when it was right just by how much it agreed with itself. It's like a student who writes out a long, detailed explanation but still gets the final answer wrong, and you can't tell the difference just by looking at how neatly they wrote it.
The "Shared Bias" Mystery
The researchers also checked if different AI families (like GPT and Claude) made the same mistakes.
- The Discovery: They found that when different AI models got a question wrong, they often picked the exact same wrong answer.
- The Analogy: It's like two different students from different schools both getting the same math problem wrong because they both learned a specific, wrong trick in class. This suggests that when AIs agree on a wrong answer, it's not a lucky coincidence; it's a shared bias baked into their training.
So, Should We Stop Using Agreement?
Not exactly. The paper doesn't say "never trust agreement." It says, "Use it carefully, and know your limits."
- Good for Budgeting: Agreement is great for deciding how much computer power to spend. If the models are all unsure (low agreement), you might want to spend more time or money to get a better answer.
- Bad for Trusting: You should never use agreement as a standalone "trust me" button. If a super-smart model says, "I am 90% sure," and it agrees with itself, you still can't be sure it's right. In fact, the paper suggests that routing hard questions to the "smartest" model based on confidence is a bad idea because that model is the most over-confident.
The Bottom Line
The paper concludes that self-consistency (agreement) is a "conditional proxy." It's a helpful hint in some situations (like with mid-level models), but it is not a guarantee of truth.
- What is proven: In these specific tests, high agreement often meant high confidence but low accuracy, especially for the most advanced models.
- What is suggested: The "over-confidence" might be a side effect of how these models are trained, and shared errors happen because different models learn from similar data.
- What is unknown: The paper notes that these results are specific to the models and questions tested. We don't know if this happens with every type of AI or every kind of question, but the warning is clear: Don't assume that just because everyone agrees, they are right. Sometimes, they are just all looking at the same wrong map.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.