← Latest papers
💬 NLP

LLM Judges Inconsistently Disagree Across Safety Criteria and Harm Categories

This paper reveals that Large Language Models used as automated judges exhibit significant inconsistency and unreliability, particularly when evaluating regulated domains like finance or varying by language and criteria, highlighting the need for careful implementation in safety assessments.

Original authors: Krishnapriya Vishnubhotla, Soumya Vajjala, Akriti Vij, Isar Nejadgholi

Published 2026-06-01
📖 5 min read🧠 Deep dive

Original authors: Krishnapriya Vishnubhotla, Soumya Vajjala, Akriti Vij, Isar Nejadgholi

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have hired a panel of six different art critics to judge a series of paintings. You want them to agree on which paintings are "safe" (good, appropriate) and which are "unsafe" (bad, harmful). You expect that if you show them the same painting, or even a slightly different version of it (like a photo of the painting in a different frame, or a translation of the artist's note into another language), they should all give roughly the same verdict.

This paper is about testing exactly that, but instead of art critics, the "judges" are Large Language Models (AI), and instead of paintings, they are judging AI-generated text for safety.

Here is the breakdown of what the researchers found, using simple analogies:

1. The "Obvious vs. Subtle" Problem

The researchers tested the AI judges on two very different types of "bad" content:

  • The "Screaming Fire" Test (Violence): They asked the AI to judge text about violence, hate speech, or illegal acts.
    • The Result: The AI judges were pretty good at this. If a painting was clearly on fire, all six critics agreed, "This is unsafe."
  • The "Financial Advice" Test (Regulated Domains): They asked the AI to judge text where an AI gives advice on things like credit scores, visa applications, or medical advice.
    • The Result: The AI judges were terrible at agreeing here. It's like asking six critics to judge a painting that is almost a masterpiece but has one tiny, subtle flaw. One critic says, "It's safe, just a little risky," while another says, "It's dangerous, throw it out!" They couldn't agree on what "safe" even meant in these complex, real-world scenarios.

2. The "Magic Mirror" Problem (Self-Consistency)

The researchers took a single AI response and showed it to the judges in different "mirrors":

  • Translation: They showed the text in English, then French, then Hindi, etc.

  • Paraphrasing: They rewrote the text to sound more formal, more casual, or to change the order of sentences, without changing the actual meaning.

  • The Result: The judges were inconsistent. If a judge said, "This English sentence is safe," they might look at the exact same sentence translated into Hindi and say, "Wait, this Hindi version is unsafe!" Or, they might look at a rephrased version and change their mind.

  • The Metaphor: Imagine a security guard who stops a person in a red shirt but lets the same person in a blue shirt pass, even though they are the exact same person. The AI judges are easily tricked by how the words are dressed up, not just what the words actually say.

3. The "Jury" Problem (Cross-Consistency)

The researchers also asked: "If we put all six AI judges in a room together, will they agree with each other?"

  • The Result: No. Even when looking at the exact same text in the exact same language, the judges often gave opposite answers.
  • The Metaphor: Imagine a jury where one person votes "Guilty," another votes "Not Guilty," and a third votes "Maybe." They are all looking at the same evidence, but they have completely different internal rulebooks for what counts as a crime. The paper found that for complex topics (like financial advice), the "jury" is often in total chaos.

4. The "Fake Agreement" Trap

The paper points out a tricky statistical illusion.

  • Sometimes, all the judges say "Safe" 99% of the time. If you just count the "Yes" votes, it looks like they are in perfect agreement (100% consistency).
  • The Catch: But if they are just saying "Safe" because they are lazy or biased, and not actually thinking about the specific details, that's not real agreement. The researchers used special math (like a "chance-corrected" score) to strip away this fake agreement.
  • The Metaphor: It's like a group of friends guessing the answer to a trivia question. If the answer is "Yes" 99% of the time, and they all guess "Yes," they look like geniuses. But if the question is actually hard and they are just guessing the most common answer, they aren't actually agreeing on the reasoning. The paper shows that when you look deeper, the AI judges are actually disagreeing a lot.

The Bottom Line

The paper concludes that while AI judges are okay at spotting obvious, loud dangers (like violence), they are unreliable and inconsistent when judging nuanced, real-world advice (like finance or law).

If you use an AI to act as a safety guard for complex topics, you can't trust it to be consistent. It might pass a dangerous answer just because the sentence structure changed slightly, or it might reject a safe answer just because it was translated into a different language. The "jury" of AI models is currently too divided to be trusted for high-stakes decisions without human oversight.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →