The Flaw of Averages: Measuring Benchmark-Level Distributional Robustness
This paper introduces "Harmony," an entropy-based metric for measuring the distributional robustness of language model benchmarks, demonstrating that low-Harmony benchmarks often misrepresent broad model competence through uneven subdomain performance and recommending that Harmony be reported alongside aggregate accuracy to ensure more representative evaluations.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
In the rapidly evolving world of artificial intelligence, researchers rely on standardized tests, known as benchmarks, to measure how well language models think and learn. These tests are the yardsticks of the field, determining which models are considered the most capable and guiding the direction of future development. However, just as a student might ace a math exam but fail a history test, a language model can perform brilliantly on one type of question while struggling with another. When these varied results are rolled up into a single average score, the details can disappear. A high overall score might suggest a model is broadly competent, even if that competence is actually concentrated in just a few specific areas, masking significant weaknesses elsewhere. This phenomenon creates a misleading picture of progress, where the summary statistic hides the true distribution of a model's abilities.
A team of researchers at Johns Hopkins University has developed a new way to look past these averages to see the full picture. They introduced a diagnostic tool called HARMONY, which measures how evenly a model's performance is spread across the different topics within a single test. Instead of simply asking how many questions a model got right, this approach asks whether the model is getting them right across the board or if its success is clustered in a narrow set of subjects. By analyzing nineteen different benchmarks across five families of language models, the researchers found that many popular tests suffer from a lack of balance. In these cases, the aggregate scores often overstate a model's general intelligence because the test is dominated by topics where the model happens to be strong, while its failures in other areas are drowned out by the average.
The researchers demonstrated that this imbalance is not just a theoretical concern but has real consequences for how we evaluate artificial intelligence. They took several benchmarks and removed the questions that were overly similar to one another, effectively rebalancing the test to include a wider variety of topics. On tests that were already well-balanced, this change had little effect on the final scores. However, on tests that were unbalanced, the scores shifted dramatically. For example, a benchmark evaluating performance in a medically consequential domain exhibited substantial, often statistically significant, shifts in aggregate accuracy once the overrepresented items were pruned away, revealing that the model's earlier high scores were not as robust as they appeared. In contrast, a general knowledge test remained stable, suggesting its original score was a more faithful reflection of the model's true capabilities.
This work suggests that the current method of ranking models by a single number is often insufficient. The researchers argue that to truly understand a model's competence, we must look at how uniformly it performs across different domains. They found that larger models do not automatically become more balanced; in some families of models, increasing the size actually made the performance more concentrated in specific areas, while in others, it became more even. This indicates that simply making models bigger does not guarantee they will be more generally capable. The study concludes that whenever a benchmark score is reported, it should be accompanied by a measure of this distributional balance. Only by acknowledging where a model is strong and where it is weak can the scientific community avoid the trap of believing that an average score represents a broad, reliable intelligence.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.