← Latest papers
🤖 machine learning

Whose Gold? Annotator-Pool Disagreement Is Large at the Item Level, and Hidden by Small Leaderboards

This paper reveals that significant disagreement between expert and crowd annotators at the item level is masked by small leaderboard sizes, creating a false sense of stability in model rankings that fails to generalize as the number of models increases.

Original authors: Anik Jha

Published 2026-08-18
📖 5 min read🧠 Deep dive

Original authors: Anik Jha

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

In the race to build smarter artificial intelligence, researchers rely on a simple but powerful idea: to know if a machine is getting better, you must ask a human to judge its work. This process, known as a preference benchmark, involves showing a human two different answers from an AI and asking which one is better. The collection of these human votes becomes the "gold standard," a trusted reference point against which all new models are measured. For years, the field has operated on a quiet assumption: that it does not matter which humans are hired to do the judging. The prevailing belief is that while individual people might make random mistakes, the group as a whole is a reliable instrument, measuring a single, stable truth about what makes an answer good. If this assumption holds, then the specific identity of the judges is just a minor detail, like the brand of pen used to write the score.

But a new investigation suggests that this detail is actually the most important part of the equation. Researchers set out to test whether the choice of human judges truly matters by running the same set of AI comparisons twice, using two completely different groups of people. One group consisted of ordinary crowdworkers, the kind of people hired by the thousands for online tasks. The other group was made up of screened domain experts, individuals with specialized knowledge in the field. The researchers took a massive dataset of thousands of AI comparisons and asked both groups to vote on the winners. They found that the two groups of humans did not agree with each other nearly as often as the field assumed. On nearly one in four items, the two groups chose different majority winners. In almost one out of every ten cases, the crowdworkers and the experts picked opposite winners, declaring the loser of one group to be the champion of the other.

Despite this massive disagreement at the level of individual items, the final result looked deceptively stable. When the researchers used these conflicting votes to rank the six AI models in the study, the final leaderboard was identical regardless of which group of humans did the judging. The order of the models did not change at all. This created a confusing picture: the judges were shouting different answers, yet the scoreboard remained silent. The researchers realized that this stability was not a sign of a robust system, but rather a lucky accident caused by the small number of models being compared. Because the models were spaced far apart in their performance, the small shifts in voting caused by changing the judges were not enough to push one model past another.

To test how fragile this stability really was, the researchers ran simulations to see what would happen if the leaderboard were larger, containing the dozens of models that are common in real-world competitions. The results were stark. While a six-model leaderboard remained unchanged, a leaderboard with just ten models would see a model displaced in eighty-six percent of cases. With twenty models, the chance of a reshuffle was nearly one hundred percent. The study concluded that the field's current practice of reporting small leaderboards is safe only because the models are few and far between. If the field were to publish a larger, more realistic arena of models, the choice of human judges would immediately scramble the rankings. The "gold standard" is not a fixed truth; it is a moving target that shifts depending on who is hired to hold the measuring tape.

The investigation did not stop at human judges. The researchers also examined how artificial intelligence judges, which are increasingly used to automate these rankings, handle the disagreement between human groups. They tested three different AI judges, including models from different companies, to see which group of humans they agreed with more. Every single AI judge aligned significantly more with the crowdworkers than with the experts. This suggests that the AI judges have learned to mimic the preferences of the crowd, likely because the data used to train them came from similar crowdsourced workers. The researchers found that an AI judge's agreement with humans is not just a measure of its intelligence, but also a reflection of the hiring decisions made years ago when the training data was collected.

The study also uncovered that the way these AI judges are prompted can change their minds more than the choice of human judges does. When the researchers changed the order in which the answers were presented, the AI judges reversed their decisions on nearly half of the items (specifically 44.6% for one model). Furthermore, when the researchers changed only the output format (asking for a JSON object versus a bare token) while keeping everything else fixed, the same judge agreed with itself on only 47.2% of items, showing a systematic shift in one direction. This sensitivity to presentation order and formatting is a much larger source of error than the difference between expert and crowd opinions. The researchers argue that the entire field needs to change how it reports these results. Simply stating that a model agrees with humans a certain percentage of the time is incomplete without specifying which humans were used. The safety of current rankings is an illusion created by small sample sizes, and the reliability of future rankings depends entirely on acknowledging that the choice of annotators is not a minor detail, but a fundamental variable that shapes the outcome.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →