Benchmark Illusion: Disagreement among LLMs and Its Scientific Consequences
This paper reveals a "benchmark illusion" where large language models with comparable accuracy exhibit significant disagreement on reasoning tasks, causing model selection to become a hidden but critical variable that can drastically alter or even reverse scientific findings in fields like education and political science.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are a chef trying to judge which of two new ovens is better. You put a standard cake in both, and they both come out perfectly golden brown. You declare them "equal" and decide you can swap them out anytime in your kitchen without changing the taste of your food.
This paper argues that in the world of Artificial Intelligence (AI), this is a dangerous illusion.
The authors, Eddie Yang and Dashun Wang, call this the "Benchmark Illusion." Here is the simple breakdown of what they found, using some everyday analogies.
1. The "Perfect Score" Trap
Scientists use standardized tests (called benchmarks) to grade AI models. If Model A and Model B both get a 90% score on a hard reasoning test, we assume they are basically the same smart brain. We think, "If they both got 90%, they must know the same things and make the same mistakes."
The Reality:
The authors found that even when two AIs get the exact same score, they are often answering different questions correctly.
- The Analogy: Imagine two students taking a math test. Both get 90%.
- Student A got the algebra questions right but failed the geometry ones.
- Student B got the geometry questions right but failed the algebra ones.
- If you only look at the final grade (90%), you think they are identical. But if you ask them to solve a specific geometry problem, Student A will fail, and Student B will succeed. They have different blind spots.
The paper found that even the "top" AI models disagree with each other on 16% to 66% of the questions. They aren't just making random noise; they have systematic, different ways of thinking.
2. Why This Matters for Science
Scientists are starting to use these AI models to read research papers, grade student essays, or analyze news articles. They treat the AI like a "plug-and-play" tool, assuming any high-scoring model will do the same job.
The Danger:
Because every AI has a different "blind spot," swapping one AI for another can completely change the results of a scientific study.
The Analogy: The Biased Judge
Imagine you are a judge trying to see if a new medicine causes rashes. You ask a human assistant to read patient notes and flag the rashes.
- Assistant A is very good but accidentally misses rashes in patients who took the new medicine.
- Assistant B is also very good but accidentally misses rashes in patients who took the old medicine.
If you use Assistant A, the new medicine looks super safe (because you missed the rashes). If you use Assistant B, the new medicine looks dangerous (because you missed the rashes in the control group).
The medicine didn't change. The data didn't change. Only the "judge" changed, and the conclusion flipped.
3. Real-World Examples from the Paper
The authors tested this with real studies:
Case 1: Grading Student Essays
They took a study where human teachers graded student essays to see if a new teaching method worked. They replaced the teachers with 8 different "smart" AIs.- Result: All the AIs agreed the new method worked. However, the size of the improvement varied wildly. One AI said the improvement was huge; another said it was tiny. The difference was 80%. If you picked the wrong AI, you might think the teaching method was a miracle or a waste of time.
Case 2: Russian News Analysis
They looked at a study about how Russian state media blames or praises officials for economic news.- Result: The original human study found that officials were praised for good news and blamed for bad news.
- The Twist: When they used different AIs, some agreed with the humans. But two other AIs concluded the exact opposite: that officials were blamed for good news and praised for bad news.
- The Lesson: The choice of AI model changed the entire political conclusion of the study.
4. The Big Takeaway
The paper warns us that "Accuracy" is not enough.
In the past, if a tool was "accurate," we trusted it. But in the age of AI, a tool can be highly accurate on average but still have a specific, hidden bias that ruins scientific research.
What should we do?
- Don't just pick the model with the highest score. Treat the choice of AI model like a major experimental decision (like choosing a specific chemical or a survey method).
- Test multiple models. Just as a scientist might run an experiment three times to be sure, they should run their analysis with three different high-performing AIs to see if the results hold up.
- New Rules for AI. We need to stop just measuring "how right" an AI is, and start measuring "how consistent" it is with other AIs and how stable its errors are.
In short: Just because two AIs get an "A" on their report card doesn't mean they will do the same work in your lab. If you aren't careful, the AI you choose could secretly decide the outcome of your research.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.