The Translation Tax Is Not a Scalar: A Counterfactual Audit of English-Source Cue Inheritance in Chinese Multilingual Benchmarks
This paper challenges the notion of a uniform "Translation Tax" in English-to-Chinese benchmarks by demonstrating that score inflation is not a scalar effect but rather a complex, item-dependent phenomenon driven by specific validity risks and model-family interactions, ultimately proposing a new naturalization protocol and reporting checklist to address these nuanced issues.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are taking a test in a language you don't speak fluently. The test questions were originally written in English, but someone translated them into your language (Chinese).
There is a common fear in the world of AI research called the "Translation Tax." The idea is that when you translate a test, the translation isn't perfect. It leaves behind little "ghosts" or "clues" from the original English. The fear is that AI models might not actually understand the Chinese; instead, they might be spotting these English ghosts and using them as cheat codes to get the right answer, inflating their scores.
This paper asks a simple question: Is this "Translation Tax" a single, fixed number we can subtract from everyone's score? Or is it a messy, complicated situation that depends on the specific question and the specific AI?
Here is the breakdown of what the authors found, using some everyday analogies:
1. The "Ghost" Hunt (The Back-Translation Test)
The researchers tried to catch the AI cheating. They took the Chinese questions, translated them back into English, and asked the AI the same questions again.
- The Analogy: Imagine you translate a recipe from English to French, then back to English. If the ingredients list changes slightly (e.g., "butter" becomes "margarine"), you know the translation lost some information.
- The Finding: The "ghosts" were there, but they were tiny. The AI's scores dropped only slightly when they saw the "back-translated" version. It's like finding a few crumbs of English left on a Chinese plate, but not enough to prove the AI is cheating on a massive scale.
2. The "Native vs. Foreign" Comparison
They compared AI scores on these translated tests against scores on tests that were originally written in Chinese (not translated from English).
- The Analogy: Imagine comparing a student's score on a translated history test against a test written by a local teacher.
- The Finding: It wasn't a simple story. Some AI models designed for Chinese actually did worse on the translated tests than on the native ones. This suggests the "Translation Tax" isn't a universal bonus; sometimes, the translation actually makes things harder or confusing for certain models.
3. The "Magic Rewrite" (The Stress Test)
This was the most clever part. The researchers took specific Chinese questions and asked an AI to "rewrite" them to sound more natural, without changing the meaning, the answer, or the options.
- The Analogy: Imagine a student is taking a test where the questions are written in a stiff, robotic style. You ask a friend to rewrite the questions to sound like a normal human conversation, but you keep the answers exactly the same. If the student suddenly gets the answers right after the rewrite, it means they were confused by the style of the translation, not the content.
- The Finding:
- For "Clean" Questions: If the translation was already good, rewriting it didn't change the AI's score.
- For "Dirty" Questions: If the translation had those "English ghosts" (high residue), the AI's score went up after the rewrite.
- The Twist: The researchers found a bug in their own computer code (a formatting error) that made it look like different AI families were behaving very differently. Once they fixed the bug, the "family differences" disappeared. The only thing that remained was that bad translations hurt performance, and fixing them helped.
The Big Conclusion
The paper argues that the "Translation Tax" is not a single number (like "subtract 5% from everyone's score").
Instead, think of it like potholes on a road:
- Some roads (some AI models) are fine.
- Some cars (some AI models) are better at handling bumps than others.
- Some potholes (badly translated questions) are deep and cause crashes.
- Some potholes are just a tiny bump.
You can't just put a "Tax" sign on the whole road. You have to look at each specific pothole (each question) and each specific car (each AI model) to see what's happening.
In short: Translated tests aren't perfect, and they do contain clues from the original English, but the effect is small, inconsistent, and depends entirely on which question and which AI you are testing. There is no single "magic number" to fix it.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.