← Latest papers
💬 NLP

Mind the Gap... or Not? How Translation Errors and Evaluation Details Skew Multilingual Results

This paper reveals that the apparent performance gap between high- and low-resource languages in multilingual LLMs is largely an artifact of translation errors and inconsistent answer extraction in the MGSM benchmark, which disappears when these data quality issues are corrected.

Original authors: Jan-Thorsten Peter, David Vilar, Tobias Domhan, Dan Malkin, Markus Freitag

Published 2026-06-23
📖 4 min read☕ Coffee break read

Original authors: Jan-Thorsten Peter, David Vilar, Tobias Domhan, Dan Malkin, Markus Freitag

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are a teacher trying to grade a math test for a class of students who speak different languages. You want to know if your students are equally good at math, regardless of whether they are speaking English, French, Swahili, or Bengali.

This paper is essentially a report card on how we've been grading these tests, and it reveals a shocking mistake: We were grading the students based on a broken test paper and a faulty grading machine.

Here is the story of what the researchers found, broken down simply:

1. The Initial (Wrong) Conclusion

The researchers started by looking at a popular math test called MGSM. This test takes 250 simple math problems (like "If you have 5 apples...") and translates them into 10 different languages.

When they ran their AI models (the "students") through this test, the results looked terrible for non-English speakers.

  • The Analogy: It was like a student who speaks French getting a 60% on a math test, while the English-speaking student got a 98%.
  • The Initial Thought: Everyone assumed the AI just wasn't "smart" at math when speaking French or Swahili. They thought there was a huge "language gap" where the AI struggled to reason in other tongues.

2. The Investigation: "Wait, the Test Paper is Broken!"

The researchers decided to look closer, like a detective inspecting the test questions. They found two major problems that were messing up the scores:

Problem A: The "Lost in Translation" Errors
The human translators who created the test made subtle mistakes.

  • The Metaphor: Imagine a math problem says, "Tuesday, I walked 6 times as far as Monday." But the German translation accidentally changed "Tuesday" to "Thursday."
  • The Result: The AI model, being very logical, looked at the question and said, "Wait, the days don't match up! I can't solve this!" So, it got the answer wrong. But the mistake wasn't the AI's math skills; it was the broken question.

Problem B: The "Faulty Grading Machine"
Even when the AI got the right answer, the computer program used to read the answer was too rigid.

  • The Metaphor: Imagine the AI writes the answer as "2.000" (using a dot for thousands, common in Europe). The grading machine, expecting English style, sees the dot and thinks, "Oh, that's a decimal point! That's 2.000... wait, no, that's 2." Or worse, it sees a comma and gets confused.
  • The Result: The AI actually solved the problem correctly, but the "grading machine" marked it as wrong because it didn't understand that different countries write numbers differently.

3. The Fix: Cleaning the Mess

The researchers did two things:

  1. Fixed the Questions: They went through the test, found the translation errors (like the "Tuesday vs. Thursday" mix-up), and corrected them.
  2. Fixed the Grader: They updated the computer program to be smarter about reading numbers. It learned that a comma in French might be a decimal point, and a dot in German might be a thousands separator.

4. The New Results: The Gap Disappears

When they ran the test again with the fixed questions and the smarter grader, the results changed completely.

  • The Analogy: It was like realizing the French student actually got 98% too, but we were just misreading their paper.
  • The Outcome: The huge performance gap between English and other languages almost vanished. For the strongest AI models, they performed just as well in French, Russian, or Swahili as they did in English.

The Big Lesson

The paper concludes that for the smartest AI models, there is no real "math gap" between languages. The models can do math in any language just fine.

The reason we thought there was a gap was because:

  1. The test questions were translated poorly.
  2. We were grading the answers with a tool that didn't understand local number formats.

The Takeaway:
Before we blame the AI for being "bad at math in other languages," we need to make sure our tests are actually fair and our grading tools are working correctly. The researchers released their corrected test so everyone else can use a fair ruler to measure the future.

What this paper does NOT say:

  • It does not say AI is perfect in all languages (smaller, weaker models still struggle a bit).
  • It does not say we should stop testing in other languages.
  • It does not claim that all benchmarks are broken, only that this specific one (MGSM) had these specific issues.

In short: The AI wasn't the problem; the test was.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →