Position: Evaluation Scores Are Perishable Knowledge Claims
This paper argues that evaluation scores are perishable knowledge claims prone to "trust inflation" when aggregated via averaging, proposing instead a conservative "weakest-link" approach that explicitly tracks a score's formality, scope, and validity window to prevent misleading rankings.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
The Scorecard Trap: Why "Average" Can Be a Lie
Imagine you are a detective trying to solve a mystery using clues from three different witnesses. One witness is a super-accurate robot that never lies but only saw the crime from far away. Another is a chatty human who saw everything up close but sometimes makes things up to sound cool. The third is a mirror that just repeats what the suspect says. If you take the "average" of their stories, you might get a neat, middle-of-the-road version of events that sounds perfectly reasonable. But here's the catch: if the human witness says the suspect was wearing a red hat (when they were actually wearing blue), that one wrong detail ruins the whole story, no matter how perfect the robot's description of the weather was. In the world of Artificial Intelligence, specifically when we try to measure how smart these computer brains are, we often make this exact mistake. We treat test scores like simple math problems where we can just add them up and divide by the number of tests to get a "final grade." But the authors of this paper argue that these scores aren't just numbers; they are claims about truth that have expiration dates, specific limits on where they apply, and different levels of reliability. If we ignore these hidden rules and just average everything together, we end up with "Trust Inflation"—a fancy way of saying we become overconfident in a system that might actually be broken in the most dangerous ways.
The Paper's Big Idea: Stop Averaging, Start Checking the Weakest Link
The paper, written by Sankalp and Shlok Gilda, argues that the current way we rank AI models is fundamentally flawed because it treats all test results as if they are equally strong and permanent. The authors call this problem Trust Inflation. It happens when we mix different types of evidence—like automated computer checks, human opinions, and AI judges—into a single "average" score. The problem is that a single weak point can destroy the whole picture, but averaging hides that weakness.
To explain this, the authors use a simple analogy: imagine a chain. The strength of the entire chain isn't determined by the average strength of all its links; it is determined by the weakest link. If you have a chain where 99 links are made of steel and one is made of paper, the whole chain will snap under the weight of a feather. Similarly, if an AI model is great at writing poetry but terrible at telling the truth (factuality), an "average" score might make it look like a decent all-rounder. But in the real world, if that AI is used to give medical advice, its inability to tell the truth is a disaster, regardless of how pretty its poetry is.
The paper suggests we stop using the "average" as our default method. Instead, we should use a "weakest-link" approach, especially for safety-critical tasks. This means if an AI fails even one important test, its overall rating should drop to reflect that failure, not be propped up by its other good scores.
The Three Rules of AI Truth
The authors propose that every AI test score should come with a "label" that tells us three specific things, much like a food package tells you what's inside, where it's from, and when it expires:
Formality (How strong is the evidence?): Not all tests are created equal. The paper creates a "tier" system.
- Tier F0 (The Weakest): This includes scores from other AI models judging the work or crowdsourced opinions. The authors suggest these can only be trusted up to a 0.70 reliability ceiling because AI judges often just like long, fancy answers even if they are wrong.
- Tier F1 & F2 (Stronger): These are structured computer checks or controlled human tests. They are more reliable, but even human tests have limits (reliability up to 0.95).
- Tier F3 (The Strongest): This is a formal mathematical proof. This is the only thing that can reach a perfect 1.00 reliability.
- The Rule: If you mix a Tier F0 score with a Tier F2 score, the whole result is capped at the F0 level (0.70). You can't make a weak test strong just by adding a strong test to it.
Scope (Where does this apply?): A test score is only true for the specific situation it was tested in. If an AI is tested on English questions, that score doesn't mean it's good at Spanish. If it's tested on general knowledge, it doesn't mean it's good at legal advice. The paper argues that scores should explicitly state their limits so we don't accidentally use them in the wrong place.
Validity Windows (When does it expire?): Test scores have a shelf life. Just like milk, they go bad. The authors point out that once an AI model has seen the test questions during its training (a problem called "contamination"), the score becomes meaningless. They suggest every score should have an expiration date. An AI-generated opinion might only be valid for a few weeks, while a formal math proof might last forever.
Real-World Evidence: The "Trust Inflation" in Action
The authors didn't just talk about theory; they built a real testing system for AI agents and found some scary bugs.
- The Silent Bug: They discovered that a mismatch between their computer code languages caused every single failure to be recorded as a success. The system was lying to them, inflating the scores without anyone noticing. This is "trust inflation" at the infrastructure level.
- The "Average" Lie: They tested this theory on a famous public leaderboard called HELM. They looked at 54 top AI models across 10 different scenarios.
- When they ranked the models by the average score, the top 5 models were one group.
- When they ranked them by the weakest-link (the lowest score in any category), the top 5 models were a completely different group.
- In fact, the overlap between the two top-5 lists was zero. The models that looked like the best "all-rounders" by averaging were actually the ones with the biggest hidden weaknesses. The models that looked "average" by the old method were actually the most reliable because they didn't have any catastrophic failures.
What This Means for the Future
The paper concludes with a call to action for the AI community. They want us to stop hiding the math behind the scores.
- Show your work: Every score should carry a label saying how strong the evidence is, what it applies to, and when it expires.
- Pick your poison: If we must combine scores, we should admit which method we are using. Are we being optimistic (using the average) or conservative (using the weakest link)? The paper argues that for safety, we should default to the conservative "weakest link" method, but at least we should admit we made that choice.
- Version control: Just like software code, the way we record test results should be versioned so we don't accidentally compare apples to oranges.
The authors admit this is a "position paper," meaning they are proposing a new way of thinking based on strong logic and engineering experience, rather than a massive new experiment that proves everything is 100% correct. They acknowledge that being too conservative might make good models look bad, but they argue it is better to be safe than to deploy a system that looks great on average but fails in a way that hurts people. By making the "epistemic status" (the truthfulness and limits) of scores transparent, we can make better decisions about which AI systems to trust.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.