← Latest papers
💻 computer science

LLM-Safety Evaluations Lack Robustness

This paper argues that current LLM safety research is hindered by significant noise and inconsistencies across the evaluation pipeline, and it proposes systematic guidelines to improve the robustness, fairness, and comparability of future attack and defense assessments.

Original authors: Tim Beyer, Sophie Xhonneux, Simon Geisler, Gauthier Gidel, Leo Schwinn, Stephan Günnemann

Published 2026-05-19
📖 5 min read🧠 Deep dive

Original authors: Tim Beyer, Sophie Xhonneux, Simon Geisler, Gauthier Gidel, Leo Schwinn, Stephan Günnemann

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to judge which of two new security guards is better at stopping intruders. You set up a test where you send a series of "tricky questions" to the guards to see if they accidentally let a bad guy in.

This paper argues that the current way we test Large Language Models (LLMs) for safety is like using a broken, inconsistent, and poorly designed test for those security guards. Because the test itself is flawed, we can't truly tell who is the better guard, and progress in making AI safer is getting stuck.

Here is a breakdown of the paper's main points using simple analogies:

1. The Test Questions Are Too Small and Repetitive (Datasets)

Imagine you are testing a guard's ability to spot a thief. Instead of showing them 1,000 different types of thieves, you only show them 50 pictures, and 40 of them look exactly the same.

  • The Problem: The paper says current safety tests use very small lists of "bad prompts" (often just 100–500). Because the list is so small, the results are full of "noise" (random luck). If you test the same guard on a slightly different list of 50 pictures, the score might change wildly, making it impossible to know if they are actually safe.
  • The "Subsampling" Issue: Sometimes researchers take a huge list of questions but only pick a tiny, random handful to test. It's like a teacher grading a student based on only three questions out of a 100-question exam. This makes it hard to compare different students fairly because everyone is taking a different, tiny version of the test.
  • The "English-Only" Blind Spot: Most tests are only in English. But if a guard only speaks English, they might fail a test in Spanish. The paper notes that safety knowledge often changes depending on the language, so testing only in English gives a false sense of security.

2. The Rules of the Game Are Confusing (Algorithms)

Imagine two guards are being tested, but one is allowed to use a flashlight, while the other is forced to work in the dark. Or, one guard is tested with a stopwatch that runs fast, and the other with one that runs slow.

  • Hidden Settings: The paper points out that tiny, invisible details in how the tests are run change the results drastically. For example, how the computer handles "spaces" between words or what kind of "chat template" is used can change a guard's success rate by 14% or more.
  • The "Target" Trap: Many attacks try to force the AI to say a specific phrase like "Sure, here is how..." to prove it broke. But if the AI is trained to say "Of course!" instead, the attack fails not because the AI is safe, but because the attacker was using the wrong "key." This makes the AI look safer than it really is.
  • Unfair Comparisons: Some researchers test their attack method with unlimited computer power, while others test with very little. Comparing them is like comparing a professional athlete running a race with a sprinter who has a head start.

3. The Judges Are Biased and Inconsistent (Evaluation)

After the guard answers the tricky questions, a "Judge" has to decide: "Did they fail?"

  • The Fragmented Jury: There is no single "Supreme Court" for AI safety. Some researchers use one AI to judge the answers, others use a different AI, and some use humans. These "judges" often disagree with each other. One judge might say a response is safe, while another says it's dangerous, even for the exact same answer.
  • The "Greedy" Mistake: Most tests force the AI to give the single most likely answer (like a robot that always picks the first thing that comes to mind). But in the real world, AI is like a person who might say different things depending on their mood or how many times you ask. By only testing the "most likely" answer, we miss the times the AI might accidentally say something dangerous when it's "thinking" differently.
  • The "Over-Refusal" Blind Spot: A safe guard shouldn't just stop bad guys; they shouldn't stop good people either. If a guard refuses to let a harmless person in because they look suspicious, that's a problem called "over-refusal." The paper says most tests ignore this. They only check if the guard stops bad guys, not if they are being too grumpy with good guys.

The "Opposing View" (Why things are the way they are)

The paper also listens to the other side. Some researchers argue:

  • Small tests are cheaper: Big tests cost a lot of money and time. Small tests let researchers try new ideas quickly.
  • Perfection is impossible: Language is messy. We will never have a perfect "Judge" that understands every nuance of human conversation. We just have to keep improving the best tools we have.
  • Progress happens anyway: Even with messy tests, the field is still moving forward. New ideas are being found even if the scoreboard isn't perfect.

The Solution: A Better Rulebook

The authors aren't saying we should stop testing. They are saying we need to fix the rulebook so everyone plays by the same rules. They suggest:

  1. Use bigger, better question lists so the results aren't just luck.
  2. Standardize the settings so everyone uses the same "flashlight" and "stopwatch."
  3. Use multiple judges and have humans double-check the results to catch biases.
  4. Test both sides: Check if the AI stops bad guys and if it lets good guys in.

In short: The paper claims that right now, we are trying to measure how safe AI is with a ruler that keeps changing its own length. Until we fix the ruler, we can't be sure if we are actually making progress.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →