← Latest papers
💬 NLP

Challenges and Recommendations for LLMs-as-a-Judge in Multilingual Settings and Low-Resource Languages

This paper analyzes the limited and often inconsistent application of LLM-as-a-Judge in multilingual and low-resource language settings within the NLP community, highlighting risks like overtrust and single-model reliance while offering recommendations for more robust evaluation practices.

Original authors: A. Seza Doğruöz, Xixian Liao, Verena Blaschke, Jakob Prange, Senyu Li, David Ifeoluwa Adelani

Published 2026-07-03
📖 5 min read🧠 Deep dive

Original authors: A. Seza Doğruöz, Xixian Liao, Verena Blaschke, Jakob Prange, Senyu Li, David Ifeoluwa Adelani

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are running a massive talent show for writers from all over the world. You have thousands of contestants speaking different languages, from English and Spanish to rare, low-resource languages like Yorùbá or Maasai.

In the past, to decide who wins, you needed a panel of human judges who spoke every single language. This was slow, expensive, and hard to organize.

Recently, a new tool arrived: The AI Judge. This is a super-smart computer program (a Large Language Model, or LLM) that can read answers, compare them, and give scores. Because it's fast, cheap, and speaks many languages, everyone started using it to replace the human judges.

The Problem:
This paper is like a detective report investigating what happens when we let this AI Judge run the show for languages it doesn't know very well. The authors looked at 33 recent research papers that tried to use AI Judges for non-English languages. They found that while the idea sounds great, the reality is messy and risky.

Here is the breakdown of their findings using simple analogies:

1. The "Tourist" Judge

Imagine the AI Judge is a tourist who has visited Paris (English) many times and knows the city perfectly. But now, they are asked to judge a cooking contest in a remote village in the Amazon (a low-resource language).

  • The Paper's Claim: The AI Judge is great at English. But in low-resource languages, it's like that tourist trying to judge a local dish they've never tasted. They might think the food is delicious just because it looks fancy, or they might misunderstand the ingredients entirely.
  • The Reality: The paper found that in many studies, researchers assumed the AI Judge was just as good at Yorùbá or Nepali as it was at English. They didn't check if the judge actually understood the language. Often, the AI was just guessing or hallucinating, but the researchers trusted it anyway.

2. The "Single Referee" Bias

In a sports game, if you only have one referee, and that referee makes a mistake, the whole game is unfair.

  • The Paper's Claim: Most of the studies the authors looked at used only one AI model (usually a famous one like GPT-4) to do all the judging. They didn't ask a second AI to double-check the work.
  • The Reality: It's like having a referee who is also a fan of one team. If that specific AI model has a bias (for example, it likes long answers over short, good ones), the whole evaluation is skewed. The paper found that 48% of the studies relied on just one judge model, and 33% used GPT as their only judge.

3. The "English-Only" Safety Net

Imagine a teacher grading a test in French. To make sure they are doing a good job, they check their grading against an answer key written in French.

  • The Paper's Claim: Many researchers claimed their AI Judge was reliable. But when the authors looked closer, they saw the "safety net" was missing. They often checked the AI's reliability using English data, then assumed it worked the same way for French, German, or Swahili.
  • The Reality: The paper found that in many cases, the AI was never actually tested against human experts in the target language. They validated the judge in English (where it's good) and then blindly applied it to other languages (where it might be terrible).

4. The "Overconfident" Crowd

There is a tendency in the research community to trust the AI too much.

  • The Paper's Claim: Because AI is fast and cheap, researchers are "over-trusting" it. They treat the AI's score as the absolute truth, even when the AI is struggling with a difficult language.
  • The Reality: The paper warns that for languages with very few speakers or data (low-resource), the AI is often the weakest it can be. Yet, because there are no human experts available to double-check, people just accept the AI's flawed scores. This creates a cycle where bad evaluations are accepted as facts.

The Authors' Recommendations (The "Rules of the Game")

To fix this, the paper suggests four simple rules for anyone using an AI Judge:

  1. Test the Judge in the Local Language: Don't assume the AI knows the language just because it speaks English. You must test if the AI actually agrees with human speakers of that specific language before you trust its scores.
  2. Keep a Human in the Loop: Even if you use an AI, you need a human to check a small sample of the work. It's like having a head coach review the referee's calls.
  3. Check if Old Tools Work Better: Sometimes, simple, old-fashioned math tools (like counting exact word matches) might be safer than a fancy AI that doesn't understand the language. Don't use the AI if a simpler tool works just as well.
  4. Respect Culture, Not Just Words: A language isn't just grammar; it's culture. An AI might understand the words of a story but miss the cultural meaning. Researchers need to check if the AI understands the context and culture of the language, not just the vocabulary.

The Bottom Line

The paper concludes that while AI Judges are a powerful tool, using them for languages they don't fully understand is dangerous. It's like letting a tourist judge a local art festival without knowing the local art history. Until we verify that these AI judges are actually competent in specific, low-resource languages, we shouldn't trust their scores as the final word. We need to stop assuming they work everywhere and start proving they do.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →