Recovered in Translation: Efficient Pipeline for Automated Translation of Benchmarks and Datasets
This paper introduces "Recovered in Translation," an automated framework that leverages test-time compute scaling strategies like Universal Self-Improvement and a novel T-RANK method to generate high-quality, semantically preserved translations of benchmarks into eight Eastern and Southern European languages, thereby overcoming the limitations of existing resources and enabling more reliable multilingual LLM evaluation.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are building a giant, global library of tests to see how smart different AI computers are. Right now, most of these tests are written in English. To test AI in other languages, someone has to translate these tests.
The Problem: The "Broken Telephone" Game
Currently, translating these tests is like playing a game of "Broken Telephone" with a broken phone.
- The Old Way: Existing translations often treat the question and the answer like two separate pieces of paper. They translate the question, then translate the answer, and just tape them back together.
- The Result: The connection breaks. In languages like Ukrainian or Greek, where words change their shape depending on their role in a sentence (grammar cases), this separation causes the answer to accidentally give away the correct choice. It's like translating a riddle where the answer is hidden in the grammar of the question itself. The AI doesn't solve the puzzle; it just cheats because the translation gave it a hint.
The Solution: "Recovered in Translation"
The authors of this paper built a new, automated factory called R-Translation to fix this mess. Instead of a human translator or a simple robot, they use advanced AI to act as a team of expert editors.
Here is how their factory works, using a creative analogy:
1. The "T-RANK" Method: The Olympic Judges
Imagine you have five different translators trying to translate a tricky sentence.
- Old Method (Best-of-N): You ask a judge to give each translator a score from 1 to 10. The problem? The judge might get tired or biased, giving the first translator a high score just because they spoke first.
- The New Method (T-RANK): Instead of scoring, the AI acts like an Olympic judge panel. It puts the five translations in a ring and makes them fight it out.
- Round 1: Translator A fights Translator B.
- Round 2: Translator B fights Translator C.
- The Twist: The AI rotates who goes first so no one gets a "home-field advantage."
- The Final Step: Once the winner is picked, the AI looks at all the translations again, finds the best parts of each one, and stitches them together into a perfect "Frankenstein" translation that has the best grammar and meaning from everyone.
2. The "USI" Method: The Self-Improving Chef
Think of this as a chef who makes five different versions of a soup.
- Instead of just picking the best one, the chef tastes all five, realizes, "Hmm, Soup #1 has the right salt, but Soup #3 has the perfect spice," and then mixes them into one Super-Soup that is better than any single attempt. This is called Universal Self-Improvement.
3. Why This Matters
The researchers tested this factory on 8 different languages (like Ukrainian, Turkish, and Greek) using famous AI tests like MMLU and Winogrande.
- The Result: Their new translations were much more accurate.
- The Impact: When they ran AI models on these new, high-quality tests, the scores changed. Some models looked smarter, others looked dumber. This proves that the old tests were lying to us because the translations were bad.
The Big Picture
Think of the old translations as a blurry photo of a landscape. You can kind of see the trees, but you can't tell if it's a forest or a park.
The authors' new framework is like sharpening the lens. Suddenly, you can see the details.
Why should you care?
If we want AI to be fair and helpful for everyone in the world, not just English speakers, we need to make sure the tests we use to grade them are perfect. If the test is broken, the grade is wrong. This paper provides the tools to fix the tests, ensuring that when an AI speaks Ukrainian or Greek, we are actually testing its intelligence, not its ability to spot a translation error.
In short: They built a smart, automated system that acts like a team of expert editors to translate AI tests perfectly, ensuring that AI models get a fair grade in languages other than English.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.