← Latest papers
💬 NLP

Lost in Translation: Do LVLM Judges Generalize Across Languages?

This paper introduces MM-JudgeBench, a large-scale multilingual benchmark for evaluating LVLM judges, revealing that current reward models exhibit significant cross-lingual performance variance and inconsistent behavior across 25 diverse languages, thereby highlighting the critical need for multilingual evaluation standards in automated model assessment.

Original authors: Md Tahmid Rahman Laskar, Mohammed Saidul Islam, Mir Tafseer Nayeem, Amran Bhuiyan, Mizanur Rahman, Shafiq Joty, Enamul Hoque, Jimmy Huang

Published 2026-04-22
📖 4 min read☕ Coffee break read

Original authors: Md Tahmid Rahman Laskar, Mohammed Saidul Islam, Mir Tafseer Nayeem, Amran Bhuiyan, Mizanur Rahman, Shafiq Joty, Enamul Hoque, Jimmy Huang

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have a team of super-smart art critics (these are the AI models, or "LVLMs") whose job is to look at pictures and decide which description of the picture is better.

For a long time, these critics were only tested on English pictures and English descriptions. They seemed like geniuses. But the researchers behind this paper asked a simple, scary question: "What happens when we speak to these critics in French, Hindi, or Kazakh? Do they still know what they're talking about, or do they just guess?"

Here is the breakdown of their findings, using some everyday analogies:

1. The New Test: "The Global Talent Show"

The researchers built a massive new test called MM-JudgeBench. Think of it as a global talent show for AI critics.

  • The Stage: They took existing tests and translated them into 25 different languages (from common ones like Spanish to low-resource ones like Kazakh).
  • The Contestants: They tested 22 different AI models, including the famous "big tech" ones (like GPT-5 and Gemini) and open-source community models (like Qwen).
  • The Task: Show the AI an image and two different descriptions. Ask the AI: "Which one is true, and why?"

2. The Shocking Results: "The Language Barrier"

The results were like watching a world-class pianist play beautifully in English, but then stumbling over the notes when asked to play in a different language.

  • The "English-Only" Illusion: Many models that looked perfect in English suddenly became confused or made up facts when switched to other languages.
  • The "Efficiency Trap": Some models were designed to be "lightweight" and fast (like a sports car). They worked great in English, but in other languages, they completely crashed. It's like a sports car that drives fine on a highway but falls apart on a dirt road.
  • The Kazakh Problem: The models struggled the most with Kazakh (a low-resource language). It's as if the critics had studied hard for a test in English but were suddenly asked to take a test in a language they had never seen before. They didn't just get a few questions wrong; they often failed completely.

3. The "Hallucination" Issue

In the English tests, the AI critics were good at spotting lies. But in other languages, they started hallucinating (making things up).

  • Example: In English, an AI might correctly say, "This dog is on a bench."
  • In French (incorrectly): The same AI might say, "This dog is on a bench that says 'WATSON BOWL' on it," even though the bench in the picture has no writing. The AI was so eager to sound smart in French that it invented details that weren't there.

4. The Winners and Losers

  • The Big Tech Giants: The most expensive, closed-source models (like GPT-5) were the most consistent. They were like polyglot diplomats who could switch languages without losing their cool.
  • The Open Source Heroes: Surprisingly, a model family called Qwen performed incredibly well, often beating much larger and more expensive models. They were like the underdog athletes who trained harder and adapted better than the favorites.
  • The "Specialized" Models: Some models built specifically to be judges (like LLaVA-Critic) actually performed worse than general models. It's like hiring a specialist who knows the rules of the game but forgets how to play it.

5. Why This Matters

The paper concludes that we cannot trust these AI judges in a global world yet.

If a company uses an AI to decide which customer service chatbot is "better" for a French-speaking user, but the AI judge only speaks English well, the company might pick the wrong chatbot. The AI judge might think a bad answer is good just because it's written in a language the judge doesn't fully understand.

The Bottom Line

The researchers are saying: "Stop testing your AI judges only in English." Just because a model is smart in one language doesn't mean it's smart in all of them. To build truly reliable AI, we need to test them in the messy, diverse reality of the whole world, not just in the comfortable bubble of English.

They also released a "training manual" (a dataset) to help these AI judges learn how to be better in other languages, hoping to fix the problem before it causes real-world headaches.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →