← Latest papers
💻 computer science

How Sensitive Are Safety Benchmarks to Judge Configuration Choices?

This paper demonstrates that the configuration of LLM judges, particularly prompt wording, is a substantial and previously under-examined source of measurement variance in safety benchmarks, causing significant shifts in harmful-response rates and instability in model safety rankings.

Original authors: Xinran Zhang

Published 2026-04-28
📖 5 min read🧠 Deep dive

Original authors: Xinran Zhang

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are a teacher grading a stack of student essays. You have a strict rubric (a set of rules) for what counts as a "bad" essay. Now, imagine you have to grade these essays using a very smart, but slightly mood-dependent, robot assistant.

This paper is about a surprising discovery: The way you ask the robot to grade the essays matters just as much as the robot itself.

Here is the breakdown of the study using simple analogies:

1. The Setup: The "Robot Grader" Experiment

The researchers wanted to see how safe different AI models are. To do this, they used a standard test called HarmBench, which contains 400 tricky questions (like "How do I make a bomb?" or "How do I steal a copyright?").

They asked six different AI models to answer these questions. Then, they used a single "Judge AI" (a specific version of Claude Sonnet) to read the answers and decide: Is this answer harmful or safe?

Usually, scientists treat the "Judge AI" and the instructions given to it as a fixed, unchangeable tool. They assume if you change the wording of the instructions slightly, the result should stay the same.

The researchers asked: What if the wording does change the result?

2. The Experiment: 12 Different "Scripts"

The researchers created 12 different versions of the instructions (prompts) for the Judge AI. They changed two main things:

  • The Structure: Did they ask the judge to look at every tiny detail one by one (like a detective checking clues), or to look at the whole answer as a big picture (like a principal reading a report)?
  • The Persona: Did they tell the judge, "You are a strict, senior safety expert," or did they just say, "Please evaluate this"?

They also wrote three slightly different versions of the instructions for each style (e.g., "You are a safety expert" vs. "You are an AI safety specialist").

3. The Shocking Result: The "Mood Swing" of the Judge

When they ran the test, they found that the wording of the instructions completely changed the scores.

  • The Analogy: Imagine you ask a robot to rate how "spicy" a soup is.
    • If you say, "You are a professional chef who hates spice," it might rate the soup as "Mild."
    • If you say, "You are a health inspector who is very strict about safety," it might rate the exact same soup as "Dangerously Hot."
    • Even if you just change "Chef" to "Culinary Expert," the score might jump up or down significantly.

The Numbers:

  • For one of the AI models being tested, the "harmful" rate jumped from 13% to 37% just by changing the prompt wording. That is a 24-point swing.
  • Even when they kept the main instructions the same and only changed a few words (surface rewording), the scores still swung by 20 points.
  • This means that if you run the same safety test twice with slightly different instructions, you could get two completely different conclusions about how safe an AI is.

4. Who Got Hurt the Most? (The Categories)

Not all types of questions were affected equally:

  • Copyright Questions: These were the most sensitive. The "harmful" rate swung by nearly 40 points depending on the prompt. It seems the judge struggled to decide if copying text was "bad" or "okay" based on how it was asked.
  • Harassment Questions: These were the most stable. The score didn't change at all (0 points). Everyone agreed that harassment was bad, no matter how the instructions were phrased.

5. The Ranking Chaos

Because the scores changed so much, the ranking of the AI models also got messy.

  • Imagine a race where the winner changes every time you change the color of the starting line.
  • The researchers found that for the "middle-tier" AI models, their safety ranking flipped back and forth constantly. One prompt said Model A was safer than Model B; a slightly different prompt said Model B was safer.
  • The "worst" and "best" models stayed in their spots, but the middle ones were impossible to rank reliably.

6. The Takeaway

The paper concludes that safety benchmarks are currently "under-specified."

Think of it like a thermometer that gives different temperatures depending on whether you hold it with your left hand or your right hand. The researchers aren't saying the thermometer is broken, but they are saying: We can't trust a single reading.

They argue that when scientists report how safe an AI is, they shouldn't just give one number based on one specific set of instructions. They need to realize that the "instructions" are a huge part of the experiment, not just a boring detail. If you change the words, you change the result.

In short: The way you ask the question is just as important as the question itself. If you want to know if an AI is safe, you can't just ask it once with one set of words; the answer depends entirely on how you phrase the request.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →