← Latest papers
💬 NLP

Dimensionality and Measurement Precision in HLE's Multiple-Choice Subset

This paper uses psychometric analysis to demonstrate that Humanity's Last Exam (HLE) primarily measures a single general reasoning factor rather than distinct domain-specific capabilities, and that its measurement precision is insufficient to effectively discriminate among frontier language models.

Original authors: Mayank Sharma, Savira Nadela, Tyler Matteson

Published 2026-07-31
📖 5 min read🧠 Deep dive

Original authors: Mayank Sharma, Savira Nadela, Tyler Matteson

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are a teacher trying to figure out which student in your class is the smartest. You give them a giant, 428-question test that covers everything from ancient history to advanced physics. If a student gets a high score, you might assume they are a genius at everything. But what if you want to know more? What if you want to know if they are a math whiz but terrible at poetry, or a science nerd who can't write a sentence? To answer that, you might look at their "sub-scores" for each subject. This is exactly how we evaluate Artificial Intelligence (AI) today. We use massive tests called "benchmarks" to see how smart our AI models are. But here's the tricky part: just because a test has sections labeled "Math" and "Chemistry" doesn't mean the AI actually has separate "Math brains" and "Chemistry brains." Maybe it just has one giant "Reasoning Brain" that helps it do well on everything. This paper is like a detective story for test designers. It asks: Are these subject labels real, distinct skills, or are they just a fancy way of slicing up a single, general ability? And, even more importantly, is the test actually good at telling the difference between the very best AI models, or does it just get fuzzy when things get really hard?

The paper in question, titled "Dimensionality and Measurement Precision in HLE's Multiple-Choice Subset," dives deep into a famous new AI test called "Humanity's Last Exam" (HLE). The researchers, a team from Stanford, took 29 different AI models and ran them through the text-only, multiple-choice version of this exam, which contains 428 tricky questions. They didn't just count the right answers; they used a special kind of math called Item Response Theory (IRT). Think of this like a super-precise ruler that doesn't just measure height, but also tells you how sharp the ruler is at different heights. They wanted to see two things: first, does the test actually measure eight different skills (like math, biology, history), or is it really just measuring one big "general reasoning" skill? Second, where on the "smartness scale" is this test actually good at making distinctions?

The results were a bit of a reality check for the AI community. The team found overwhelming evidence that HLE is essentially measuring just one thing: a general reasoning ability. They calculated a statistic called McDonald's ωh\omega_h, which came out to a staggering 0.998. To put that in perspective, if 1.0 means "perfectly one single skill," this test is almost perfectly one-dimensional. The labels we put on the questions—like "Math" or "Engineering"—explained only 3.5% of the differences in how models answered. It's like if you took a group of people and tried to sort them by "running speed" and "swimming speed," but you found out that everyone who was good at running was also good at swimming, and the "swimming" label didn't tell you anything new that "running" didn't already tell you. The researchers also found that the ability estimates for specific domains were nearly identical to the total score (correlations of r ≥ 0.81), meaning the sub-scores are basically redundant.

Furthermore, the paper discovered a problem with where the test is most accurate. The test is very good at telling the difference between average AI models and slightly better ones. However, the "precision" of the test drops off a cliff for the very smartest models. The test is most precise around an ability level of θ=0.35\theta = -0.35, but the top-tier "frontier" models sit at θ>0\theta > 0. In this high-ability zone, the test becomes blurry. It's like a thermometer that is great at telling the difference between a cold day and a mild day, but when it gets to "boiling hot," the needle just sticks and can't tell you if it's 212°F or 250°F. Interestingly, the "Engineering" section of the test was the only one that kept some of its sharpness for these super-smart models (concentrating 69.1% of its information at θ>0\theta > 0), but since Engineering is only 4% of the whole test, it's not enough to save the day.

So, what does this mean? The authors suggest that we shouldn't treat the different subject scores on HLE as proof that an AI has distinct, separate skills in math or chemistry. Instead, we should probably view the test as a measure of general reasoning. More importantly, as AI models keep getting smarter, this test might stop being useful for telling the very best models apart, because it wasn't designed with enough "hard" questions to distinguish between the absolute giants. The paper doesn't say the test is useless, but it does warn us that we need to be careful not to over-interpret the sub-scores and that we might need new, even harder tests to keep up with the rapid progress of AI.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →