← Latest papers
💬 NLP

GAIA-v2-LILT: Multilingual Adaptation of Agent Benchmark beyond Translation

This paper introduces GAIA-v2-LILT, a rigorously re-audited multilingual extension of the GAIA benchmark that employs a refined workflow combining functional alignment, cultural adaptation, and difficulty calibration to demonstrate that minimal machine translation significantly distorts agent performance metrics and that proper localization substantially narrows the multilingual performance gap.

Original authors: Yunsu Kim, Kaden Uhlig, Joern Wuebker

Published 2026-04-29
📖 4 min read☕ Coffee break read

Original authors: Yunsu Kim, Kaden Uhlig, Joern Wuebker

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to test how good a new, super-smart robot assistant is at solving complex puzzles. You have a perfect set of puzzles written in English. But now, you want to see if the robot can solve these same puzzles in Arabic, German, or Korean.

The common way to do this is to take the English puzzles, run them through a standard translation app, and hand the result to the robot. The authors of this paper argue that this is like giving the robot a map that has been translated by a machine that doesn't understand the terrain. The map might look like a map, but the street names are wrong, the directions are confusing, and the landmarks don't exist.

Here is what the paper actually found and did, explained simply:

The Problem: The "Broken Map"

When researchers just use machine translation to turn English agent benchmarks (puzzles for AI) into other languages, two big things go wrong:

  1. The "Literal Trap" (Functional Misalignment): Sometimes the translation app gets confused by specific instructions.

    • The Analogy: Imagine a puzzle asks for the "IOC code for Cuba," which is CUB. A machine translator might see "CUB" and think, "Oh, that sounds like 'cub' as in a baby lion!" and translate the answer to the Arabic word for "lion cub."
    • The Result: The robot tries to find a lion cub, fails, and gets the puzzle wrong. It wasn't the robot's fault; the instructions were broken.
  2. The "Lost in Translation" Trap (Cultural Misalignment): Sometimes the puzzle relies on things that only exist in the US or UK.

    • The Analogy: Imagine a puzzle asks about recycling water bottles on a road trip across the US using "miles" and "dollars." If you just translate the words into Korean, the robot is now asked to drive from California to Maine (places it can't find on a Korean map) and calculate points in a currency that doesn't exist in Korea.
    • The Result: The robot gets stuck trying to solve a puzzle that makes no sense in its own world.

The Solution: The "Expert Editor" Workflow

Instead of just hitting "translate," the authors created a special workflow called GAIA-v2-LILT. Think of this as hiring a team of expert editors to fix the map before giving it to the robot.

Their process has three layers:

  1. The Robot Check (Automated): A computer script quickly checks for obvious errors, like "Did the translator accidentally leave the answer in English?" or "Did it leak the answer inside the question?"
  2. The AI Judge (LLM Review): They use other AI models to check if the puzzle still makes sense logically and if the tone sounds natural.
  3. The Human Expert (Bilingual Linguists): This is the most important part. Real humans who speak both languages review the puzzles. They fix the "lion cub" errors, change the US road trip to a local Korean road trip, and make sure the difficulty level is the same as the original English version.

The Results: Fixing the Map Fixes the Score

The authors tested this on five languages (Arabic, German, Hindi, Korean, and Portuguese).

  • Before the fix: When they used the raw, machine-translated puzzles, the robots scored very low. It looked like the robots were terrible at these languages.
  • After the fix: Once the human experts cleaned up the puzzles, the robots' scores jumped up dramatically. In some cases, their success rate improved by 32.7%.

The Big Takeaway:
The paper concludes that the robots weren't actually "bad" at these languages. The tests were just broken. A huge chunk of the gap between how well robots do in English versus other languages isn't because the robots are less smart; it's because the puzzles were poorly translated.

By fixing the puzzles with this careful "human-in-the-loop" process, the robots' performance in other languages got much closer to their English performance (within about 3% in the best cases). This proves that if you want to know how smart an AI really is, you have to make sure the test itself is fair and culturally correct, not just a rough translation.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →