Quantifying the Impact of Translation Errors on Multilingual LLM Evaluation
This paper investigates the reliability of multilingual LLM evaluations by demonstrating that translation errors in machine-translated benchmarks significantly contribute to accuracy drops and that automatic error detection methods show non-trivial agreement with human annotations.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are a teacher trying to grade a class of students who speak different languages. You have a famous test written in English (the "Gold Standard"). To see how well your students speak Spanish, French, or German, you take that English test, run it through a translation machine, and give the translated versions to your students.
The big question this paper asks is: What happens if the translation machine makes mistakes?
The authors, a team of researchers, discovered that these translation errors aren't just tiny typos; they are like potholes in the road that can cause even the smartest students to crash, even if they knew the answer perfectly in English.
Here is a breakdown of their findings using simple analogies:
1. The "Robot Grader" vs. The Human Expert
The researchers wanted to know: Can we trust other AI models to spot these translation potholes, or do we need a human expert?
- The Experiment: They took a set of translated test questions and asked four different "Robot Graders" (AI models like GPT-5.2, Mistral, etc.) to highlight exactly where the translation went wrong. They compared the robots' work to a "Gold Standard" created by professional human linguists.
- The Result: The robots were okay, but not perfect. One robot, GPT-5.2, was the best at finding the potholes, agreeing with the human experts about half the time. The other robots were much less accurate.
- The Catch: The robots often got confused about where the error started and ended. It's like two people trying to draw a circle around a stain on a shirt; they might agree there is a stain, but one draws a tiny circle around the center, while the other draws a huge circle including the whole sleeve.
2. The "Source of the Problem" (Who is to blame?)
Sometimes, the English test question itself is weird or confusing. Sometimes, the translation makes it worse. The researchers wanted to know: Is the student failing because the English question was bad, or because the translation messed it up?
- The Analogy: Imagine a math problem that says "If a car travels 50 miles in 2 hours..."
- Source Issue: The English question is broken and says "If a car travels 50 miles in 2 apples..." (The original is nonsense).
- Translation Issue: The English is fine, but the translator changes "miles" to "kilometers" and "hours" to "minutes," making the math impossible to solve.
- The Finding: The researchers found that translation errors are the real troublemakers. Even when the English question was perfect and solvable, the translated version caused the AI students to get the answer wrong about 6 to 11 percentage points more often just because of the translation mistakes.
3. The "Score Drop" vs. The "Ranking"
This is the most surprising part. The researchers asked: If we magically fixed all the translation errors, would the leaderboard of AI models change?
- The Analogy: Imagine a race where every runner starts 10 meters behind the starting line because of a translation error.
- The Result: If you move everyone forward 10 meters (fix the translation), everyone gets faster. The runner who was in 1st place is still in 1st place. The runner in 10th place is still in 10th place. The order of the winners doesn't change.
- The Problem: However, the times are all wrong. You think the winner ran a 10-second race, but they actually ran a 9-second race. You are underestimating everyone's true ability.
- The Conclusion: Translation errors act like a uniform weight attached to every AI model. They drag everyone's score down equally. So, while the "Leaderboard" (who is #1) stays the same, the actual scores are biased and too low. We are telling the world, "This AI is only 80% smart," when it might actually be 85% smart, just because the test was translated poorly.
Summary of the Takeaway
The paper concludes that using machine-translated benchmarks to test AI is risky.
- Robots aren't perfect at spotting translation errors yet.
- Translation errors cause real, measurable drops in performance (about 6–11 points lower accuracy).
- The rankings of AI models are stable (the winner is still the winner), but the scores are unfairly low.
It's like judging a cooking competition where the ingredients were translated from a different language. The best chef might still win, but the judges might think the food tastes worse than it actually does because the recipe was slightly garbled in translation.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.