← Latest papers
🤖 machine learning

Improving Reproducibility in Evaluation through Multi-Level Annotator Modeling

This paper addresses the reproducibility crisis in generative AI evaluation by introducing a multi-level bootstrapping approach that leverages large-scale datasets with persistent rater identifiers to model annotator variance and determine the optimal balance between the number of items and responses per item required for statistical significance.

Original authors: Deepak Pandita, Flip Korn, Chris Welty, Christopher M. Homan

Published 2026-05-14
📖 5 min read🧠 Deep dive

Original authors: Deepak Pandita, Flip Korn, Chris Welty, Christopher M. Homan

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to judge which of two new AI chatbots is "safer" and more helpful. You ask a group of people to rate them. But here's the problem: if you only ask three people to rate ten different conversations, you might get a lucky (or unlucky) result that doesn't reflect reality. This is the "reproducibility crisis" the paper talks about—science is struggling to get consistent results because we aren't measuring things carefully enough.

The authors of this paper argue that we are making a big mistake by treating every human opinion as if it came from a completely different person who has no connection to the others. In reality, the same person often rates multiple items, and they carry their own personal "flavor" or bias with them.

Here is a breakdown of their findings using simple analogies:

1. The "Taste Test" Problem

Imagine you are hosting a massive taste test for two new flavors of ice cream (Model A and Model B).

  • The Old Way (What most people do): You ask 100 different people to taste one scoop each. You assume everyone's opinion is independent.
  • The Reality: In many studies, you actually ask just 5 people to taste 20 scoops each.
  • The Issue: If "Bob" happens to hate chocolate, he will rate every chocolate scoop poorly, regardless of the brand. If you treat Bob's 20 opinions as 20 independent data points, you trick yourself into thinking you have a huge amount of evidence. You are actually just listening to Bob's voice 20 times.

The paper shows that when you ignore the fact that the same people are rating multiple items, you get false confidence. You think you have found a clear winner between the two AI models when, in reality, you just haven't asked enough different people to get a true average.

2. The "Multi-Level" Solution

The authors propose a new way of looking at the data, which they call Multi-Level Bootstrapping.

  • The Analogy: Instead of just counting votes, imagine you are a detective who realizes that "Bob" is a grumpy man who always gives low scores. You need to account for Bob's grumpiness when you calculate the final score.
  • How they did it: They took real-world datasets where they knew exactly who rated what (like a dataset of 350 conversations rated by 123 people). They used a computer simulation to "resample" the data, effectively asking: "What if we had asked different people? What if we asked the same people again?"

3. The Big Discovery: More People, Not Just More Items

The paper answers a critical question: Should we ask more people to rate fewer items, or fewer people to rate more items?

  • The Finding: It is much better to have more people (raters) rating fewer items than to have a few people rating everything.
  • The Metaphor: Imagine trying to guess the average height of a city.
    • Strategy A: Ask 1 person to measure 1,000 different buildings. (This is bad; if that person is short-sighted or uses a broken tape measure, your whole city's data is wrong).
    • Strategy B: Ask 1,000 different people to measure 1 building each. (This is better; even if one person makes a mistake, the average of 1,000 people will be accurate).

The paper found that to get a statistically reliable result (a "significant" finding), you often need to increase the number of people (K) significantly, sometimes up to 100 people per item, rather than just adding more items to the list.

4. The "Batch" Effect

The researchers also looked at how data is collected in "batches" (groups).

  • The Analogy: Imagine a classroom where the teacher hands out a test to the whole class at once. The students might chat and influence each other, or they might all be tired at the same time.
  • The Finding: When people rate items in groups (batches), their opinions become even more correlated. The paper shows that if you ignore these "batches," you underestimate how much data you actually need. You might think you need 500 ratings to be sure, but because of the "batch" effect, you might actually need 2,500.

5. The Bottom Line

The paper concludes that the current standard of using only 3 to 5 raters per item is often not enough to trust the results, especially when trying to prove that one AI model is safer than another.

  • The Recommendation: If you want to know if an AI is truly safe or better, you need to stop treating human raters as interchangeable coins. You need to acknowledge that every human has a unique "voice." To get a true answer, you need to hire more humans to give their opinions, even if it costs more money.

In short: Don't just ask the same few people to judge a thousand things. Ask a thousand different people to judge a few things. It's the only way to stop fooling yourself with fake statistics.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →