← Latest papers
💬 NLP

Contrastive ESA: Human Evaluation of Multiple Translations at Once

The paper introduces Contrastive Error Span Annotation (cESA), a human evaluation protocol that assesses multiple translations simultaneously to reduce annotator noise and time while producing absolute quality scores for interpretable model rankings.

Original authors: Vilém Zouhar, Roman Grundkiewicz, Sara Rajaee, Parker Riley, Martin Popel, Rachel Bawden, Philipp Koehn, Marine Carpuat, Tom Kocmi

Published 2026-07-30
📖 7 min read🧠 Deep dive

Original authors: Vilém Zouhar, Roman Grundkiewicz, Sara Rajaee, Parker Riley, Martin Popel, Rachel Bawden, Philipp Koehn, Marine Carpuat, Tom Kocmi

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are a judge at a massive talent show, but instead of singers, you are judging translations of stories from one language to another. For years, the standard way to do this was to hand the judge one single translation, ask them to grade it in isolation, and then move to the next. It's like asking someone to rate a single apple without seeing any other apples to compare it to. The problem? It's exhausting, expensive, and surprisingly messy. When you look at just one thing alone, your brain gets tired, your mood changes, and your grades can be all over the place. This is the world of "Machine Translation Evaluation," a field dedicated to figuring out how well computers can translate human languages. The goal is simple: find the best AI translator so we can trust it with real-world tasks. But to do that, humans have to act as the referees, and the old rules of refereeing were making the game slow and the scores unreliable.

Enter a team of researchers who decided to shake up the rules of the game. They introduced a new protocol called Contrastive Error Span Annotation (cESA). Instead of showing the judge one translation at a time, they put three (or more) translations side-by-side on the same screen. It's like giving the judge a basket of apples instead of a single fruit. Suddenly, spotting a bruised apple becomes much easier because you can see the shiny, perfect ones right next to it. The researchers found that this "group viewing" method didn't just make the judges faster; it made them more consistent and accurate. They discovered that showing three translations at once was the sweet spot, saving about 31% of the time while actually improving the quality of the scores. They also checked to make sure that seeing the other translations didn't trick the judges into being unfair, and found that the bias was so tiny it didn't matter. In short, they proved that when you let judges compare options together, everyone gets a better, faster, and fairer result.

The Old Way: The Lonely Apple

For a long time, the standard way to test AI translators was "pointwise evaluation." Imagine you are grading a stack of essays. In the old method, you would pick up one essay, read it, give it a score, put it down, and then pick up the next one without ever looking back. You might give the first essay an 85 because it seemed good, but then you might give the second one a 90, even if it was actually worse, just because you were tired or the first one was really bad. This is called "annotation noise." It's like trying to judge the temperature of a cup of coffee by touching it once, in a cold room, without a thermometer. The paper notes that this method suffers from high "annotator noise" and is very costly because you need so many judges to get a reliable average.

The New Way: The Fruit Basket

The authors, led by Vilém Zouhar and colleagues, proposed a new method called cESA. Instead of looking at one translation in a vacuum, the annotator sees multiple translations of the same source text at the same time. Think of it as a "fruit basket" approach. If you see a slightly bruised apple next to a perfect one, you notice the bruise much faster than if you were looking at the bruised apple alone.

In this new system, the judge looks at a source text (like a sentence from a story) and sees, say, three different AI models' translations of it side-by-side. The judge then:

  1. Marks specific "error spans" (the exact words that are wrong).
  2. Decides if the error is "major" (a big mistake) or "minor" (a small slip).
  3. Gives a score from 0% to 100% based on a clear scale.

The paper explains that this works because of a psychological concept called "joint vs. separate evaluation." When you see things together, you can spot mistakes in one translation by comparing them to the others. It also helps you understand the "space of possible translations," making your judgment more objective. Plus, you don't have to re-read the original source text over and over again for every single translation, which saves a huge amount of time.

The Experiment: Testing the Fruit Basket

To see if this new method actually worked, the researchers ran a massive experiment using English-to-Japanese translations. They tested 12 different AI models (including big names like Gemini, Claude, and DeepSeek) using a dataset of 97 segments from news, social media, and videos.

They compared the old way (showing one translation at a time) with the new way (showing 1, 2, 3, or 4 translations at a time). Here is what they found:

  • Speed: The new method was significantly faster. When showing three models at once (the "sweet spot"), the average time to annotate a segment dropped from 77.8 seconds to 53.7 seconds. That is a 31% reduction in time per item.
  • Quality: The scores were more stable. The researchers measured "inter-annotator agreement" (how much two different judges agreed) and "intra-annotator agreement" (how much the same judge agreed with themselves on a second try). The new method showed higher agreement, meaning the judges were more consistent.
  • The Magic Number: They tested showing 1, 2, 3, and 4 models at once. They found that showing 3 models was the best balance. Going from 1 to 2 models saved a lot of time, but adding a 4th model didn't save much more time and actually made the screen a bit too crowded. Three was the perfect number to keep the judges happy and efficient.

Addressing the "What Ifs": Are the Judges Biased?

The researchers were worried that maybe seeing the other translations would trick the judges. They asked: "If a really bad translation is sitting next to a good one, will the good one look too good? Or if a great one is next to a bad one, will the bad one look too bad?"

They ran tests to check for two types of bias:

  1. Position Bias: Does the translation on the left get a better score than the one on the right? The data showed almost no difference based on position.
  2. Surrounding Bias: Does seeing a "decoy" (a very bad or very good translation) change the score of the others? They found a tiny effect: if a model was next to a much stronger model, its score dropped by about 0.5 points on average. However, the authors note this is so small compared to the actual differences between models (which can be 10 or 20 points) that it doesn't really matter. It's like if a runner finishes 10 seconds behind the winner, it doesn't matter if they finished 0.5 seconds slower because they were standing next to a world champion.

The Verdict

The paper concludes that cESA is a better way to judge AI translators. It is faster, cheaper, and produces more reliable scores. The authors suggest that anyone doing human evaluation of translations should switch to showing three models at a time. They even made a simple tool (called Pearmut) that anyone can use to run this new protocol, complete with tutorials and guidelines.

In the end, the paper suggests that by letting judges compare apples in a basket rather than one by one, we get a much clearer picture of which AI is truly the best. It's a simple change in how we look at the work, but it makes the whole process of judging machines much smarter and more human-friendly.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →