HLE-Verified: A Systematic Verification and Structured Revision of Humanity's Last Exam
This paper introduces HLE-Verified, a systematically revised and certified version of the Humanity's Last Exam benchmark that employs a rigorous two-stage expert validation and repair workflow to eliminate noisy items, thereby significantly improving the accuracy and reliability of frontier language model evaluations.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are a teacher preparing the "Final Exam of Humanity" for a class of super-smart AI students. This exam, called HLE (Humanity's Last Exam), is supposed to be the ultimate test to see who is the smartest AI in the world. It covers everything from math and physics to history and coding.
However, after the exam was released, the community started noticing something weird: the exam itself was broken.
Some questions were written in confusing riddles. Some had the wrong answers printed in the back of the book. Some had explanations that didn't make sense. It was like a math test where the question asked for "2 + 2," but the answer key said "5," and the teacher's notes tried to explain why 2+2 equals 5 using magic.
When the AI students took this broken test, their scores didn't actually reflect how smart they were. Instead, their scores reflected how good they were at guessing what the confused teacher meant to ask, or how lucky they were to stumble on the right answer despite the wrong key.
The Solution: HLE-Verified
The authors of this paper (a team from Alibaba and Qwen) decided to fix the mess. They created HLE-Verified.
Think of this not as a new test, but as a rigorous "Quality Control" audit of the old one. They didn't just guess; they built a factory line to inspect every single question.
How They Fixed It (The Two-Stage Factory)
Stage 1: The "Gold Standard" Check
They took every question and asked two types of inspectors:
- Human Experts: Real scientists and professors who know the subject inside out.
- AI Inspectors: Other powerful AIs trying to solve the problem to see if they get the same answer as the key.
If a question was clear, the answer was right, and the explanation made sense, it got a "Gold Star" and stayed exactly as it was. (668 questions survived this stage).
Stage 2: The "Surgery" Room
For the questions that were broken but fixable, they didn't throw them away. Instead, they performed surgery.
- If the question was missing a crucial detail, they added it.
- If the answer was wrong (e.g., a sign error in physics), they corrected it.
- If the explanation was nonsense, they rewrote it.
Crucially, they promised not to change what the question was trying to test. They just fixed the broken parts so the test actually measured the AI's brain, not its ability to decode a typo. (1,143 questions were fixed this way).
The "Uncertain" Box
For the remaining questions (689 of them) that were so confusing or complex that even the experts couldn't agree on the right answer, they didn't delete them. Instead, they put them in a special "Uncertain Box." They labeled them clearly: "We don't know if this is right yet; please, future experts, come help us figure this out."
What Happened When They Re-Tested?
The team took 8 of the world's smartest AI models and gave them the Original Broken Exam and the New Fixed Exam.
Here is what they found:
The Scores Went Up (A Lot):
On the original broken exam, the AIs struggled because the questions were traps. On the fixed exam, their scores jumped by 7% to 10% on average.- Analogy: Imagine a runner tripping over a hidden hole in the track. Once you fill the hole, the runner doesn't get faster legs, but they finally run at their true speed.
The "Broken" Questions Were the Worst Offenders:
For the specific questions that had wrong answers or confusing text, the improvement was massive. The AIs' accuracy jumped by 30% to 40% on these items.- Analogy: It turns out the AIs weren't "bad at math"; they were just trying to solve a math problem where the numbers were written in invisible ink. Once the ink was revealed, they solved it perfectly.
Confidence Became Real:
Before, when an AI faced a broken question, it would often guess confidently but get it wrong (because the question was a trick). After the fix, the AI's confidence matched its actual ability. If it knew the answer, it was confident. If it didn't, it wasn't.
Why Does This Matter?
This paper teaches us a vital lesson: You can't measure how smart someone is if the test is broken.
- For AI Developers: It shows that when an AI fails a hard test, it might not be the AI's fault; it might be the test's fault.
- For Science: It proves that we need to constantly check our "rulers" to make sure they aren't bent.
- For Everyone: It's a reminder that in a world of complex information, verification is just as important as creation.
In short, HLE-Verified is the "spell-check" and "fact-check" for the most important test in the history of AI, ensuring that when we say an AI is "smart," we actually mean it.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.