← Latest papers
💬 NLP

Multilingual Prompt Localization for Agent-as-a-Judge: Language and Backbone Sensitivity in Requirement-Level Evaluation

This paper demonstrates that the performance rankings of developer-agent frameworks are highly sensitive to the evaluation language and judge backbone, revealing that no single model dominates across all languages and that localizing judge instructions is critical for accurate multilingual assessment.

Original authors: Alhasan Mahmood, Samir Abdaljalil, Hasan Kurban

Published 2026-04-07
📖 4 min read☕ Coffee break read

Original authors: Alhasan Mahmood, Samir Abdaljalil, Hasan Kurban

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are hiring a team of expert architects (the AI coding agents) to build a house. You want to know if they did a good job. Usually, you'd send in a single, strict inspector (the AI Judge) who only speaks English to check their work.

This paper asks a simple but revolutionary question: What if the inspector speaks a different language than the blueprints?

The researchers discovered that the language the inspector speaks doesn't just change how they talk; it completely changes who they think is the best architect.

Here is the breakdown of their findings using everyday analogies:

1. The "Accent" Problem

Imagine you have two famous chefs, Chef GPT and Chef Gemini.

  • When a judge who speaks English tastes their food, Chef GPT gets a 5-star rating, and Chef Gemini gets a 3-star.
  • But, when a judge who speaks Arabic or Hindi tastes the exact same food, the results flip! Chef Gemini suddenly gets 5 stars, and Chef GPT drops to a 3.

The Lesson: The "best" AI isn't a universal truth. It depends entirely on the language the person grading them is speaking. If you only test in English, you might be promoting the wrong AI for the rest of the world.

2. The "Translator vs. The Brain" Experiment

The researchers ran a fascinating experiment to see why this happens. They took the same task and the same AI judge, but they changed the instructions:

  • Scenario A (Full Localization): The judge gets the task description in Hindi, and the instructions on how to grade it are also in Hindi.
  • Scenario B (Partial Localization): The task description is in Hindi, but the instructions on how to grade it are still in English.

The Result: In Scenario B, the judge's performance crashed. It was like asking a brilliant chef to cook a complex Indian dish but giving them the recipe in English while they only speak Hindi. They got confused and failed.

  • Takeaway: You can't just translate the homework; you have to translate the teacher's instructions too. If you don't, the AI gets lost in translation.

3. The "Weak vs. Strong" Student

The study looked at six different AI "judges" (backbones).

  • The Top Students (GPT-4o and Gemini): These are smart enough to understand nuance. Because they are smart, they react differently to different languages. One might be great at English logic but struggle with Arabic nuances, while the other is the opposite. Their rankings flip-flop depending on the language.
  • The Struggling Students (DeepSeek and Qwen): These models were so confused by the coding tasks that they failed almost everything, no matter what language was used. Because they were failing so hard, the language didn't matter much to them—they were just "bad" in every language.
  • The Lesson: You only see these language quirks when the AI is actually smart enough to have an opinion.

4. The "Agreement" Problem

The researchers asked all six judges to grade the same 55 coding projects.

  • The Result: They barely agreed with each other. It was like asking six different art critics to grade a painting, and they all gave it completely different scores.
  • The Metaphor: If you ask a panel of judges to rate a movie, and one loves it while another hates it, you can't trust the "average" score. The paper shows that in AI coding, the "score" is often just a reflection of which specific AI model happened to be the judge that day.

Why Does This Matter?

Currently, most AI benchmarks are like a monolingual sports league. They only test in English.

  • The Risk: If you build an AI app for a company in India or the Middle East, and you only tested it with an English-speaking AI judge, you might pick the wrong tool. You might think "Model A" is the best because it won the English tournament, but "Model B" might actually be the champion in Hindi or Arabic.

The Bottom Line

The paper is a wake-up call: Language is not just a setting; it's a variable.

If you want to know if an AI is good at coding, you can't just ask it in English. You have to ask it in the language it will actually be used in, and you have to make sure the "teacher" grading it speaks that language fluently too. Otherwise, you're just playing a game of chance with your technology.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →