← Latest papers
💬 NLP

Jury Duty: Calibration and Orientation Failures in MLLM-as-a-Judge Under Cultural Ambiguity

This paper introduces the VOIR DIRE benchmark to demonstrate that MLLM-as-a-Judge systems suffer from distinct calibration and orientation failures when evaluating culturally ambiguous content, revealing that model biases toward specific cultural norms persist even after correcting for scale compression and cannot be fully resolved by persona prompting or in-context demonstrations.

Original authors: Daniel Lee, Harsh Sharma, Eunkyu Park, Pranav Narayanan Venkit, Jeonghwan Kim, Kah Mun Chia, Andreas Vlachos, Shafiq Joty

Published 2026-06-23
📖 5 min read🧠 Deep dive

Original authors: Daniel Lee, Harsh Sharma, Eunkyu Park, Pranav Narayanan Venkit, Jeonghwan Kim, Kah Mun Chia, Andreas Vlachos, Shafiq Joty

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you hire a super-smart robot judge to grade student essays or rate photos. Usually, we check if this robot agrees with human teachers to see if it's doing a good job. But this paper asks a tricky question: Which human teachers is the robot agreeing with?

The authors, Daniel Lee and his team, built a special test called VOIR DIRE (named after the legal process of questioning jurors to find hidden biases). They wanted to see if these AI judges have a hidden "cultural bias" when they look at pictures, specifically comparing how they see things through an American lens versus a Mainland Chinese lens.

Here is the story of what they found, explained simply:

1. The Setup: Two Different Worlds

The researchers created 626 pairs of images. Each pair showed the same "concept" (like a celebratory meal, a fashion outfit, or a building) but styled in two different ways: one very American, one very Chinese.

They asked two groups of human experts to rate these images:

  • Group A: American experts.
  • Group B: Chinese experts.

The Twist: The experts agreed with each other within their own groups (they were consistent), but they disagreed strongly with each other on the same images.

  • Example: An image of a convenience store meal might be rated "low quality" by Americans but "fresh and reliable" by Chinese experts. Both groups are "right" based on their own cultural rules, but the numbers are totally different.

2. The Problem: The Robot Judge is "Stuck"

When the AI judges looked at these images, they didn't act like a neutral referee. They made two specific mistakes:

Mistake A: The "Positivity Floor" (The Rubber Band)

Imagine a ruler where the bottom numbers (1 and 2) represent "bad" or "low quality."

  • The Human Reality: Americans often rate certain Chinese cultural items as "bad" (giving them 1s or 2s).
  • The Robot's Problem: The AI judges refuse to use the bottom of the ruler. They almost never give a 1 or a 2. They get stuck at a "floor" of about 3.
  • The Result: Because the AI won't give low scores, it accidentally agrees with the Chinese experts (who rated those items higher) and disagrees with the American experts (who rated them low). The AI isn't "choosing" China; it's just too polite to say "this is bad."

Mistake B: The "Default Setting" (The Compass)

Even when the researchers tried to fix the "politeness" issue by telling the AI, "Pretend you are an American," the AI didn't fully switch gears.

  • It's like telling a GPS, "Drive to New York," but the car keeps drifting slightly toward Chicago because its internal compass is stuck pointing North.
  • The AI has a default cultural orientation. Even when told to be American, it still leans slightly toward Chinese cultural norms when judging Chinese images. This "drift" couldn't be fixed by just changing the instructions.

3. The Experiments: Trying to Fix the Bias

The team tried various tricks to see if they could make the AI a fairer judge:

  • Changing the Persona: They told the AI, "You are a typical American" or "You are a typical Chinese person."
    • Result: It helped a little, but the AI still didn't fully switch. It couldn't learn to give low scores to things Americans dislike, even when told to be American.
  • Showing Examples (In-Context Learning): They showed the AI examples of how humans rated similar pictures before asking it to judge.
    • Result: This actually made things worse. Instead of learning to use the full scale (1 to 5), the AI just started giving even higher scores (4s and 5s). It didn't learn to be fair; it just learned to be more enthusiastic.

4. The Big Takeaway

The paper concludes that you cannot just look at one "agreement score" to see if an AI judge is good.

  • The Old Way: "This AI agrees with humans 80% of the time." (But which humans? Americans? Chinese? Both?)
  • The New Way: You have to report two scores: "How well does it agree with Americans?" and "How well does it agree with Chinese people?"

If an AI agrees with Americans but disagrees with Chinese people (or vice versa), that isn't a mistake in the data—it's a property of the AI. It reveals which cultural "lens" the AI is wearing.

The Bottom Line

AI judges are like jurors who haven't been fully screened. They have hidden cultural biases that make them "agree" with one group of people while silently disagreeing with another. The paper argues that we need to stop pretending these judges are neutral and start measuring exactly whose culture they are representing.

Key Metaphor:
Think of the AI as a thermometer that is stuck at "Room Temperature."

  • If you put it in a freezer (a culture that rates things low), it won't drop to "Freezing." It stays at "Room Temp."
  • If you put it in an oven (a culture that rates things high), it might go up a bit, but it won't get as hot as it should.
  • The paper says: "Don't just say the thermometer is 'accurate' because it matches the oven. Tell us it's broken because it refuses to measure the cold."

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →