← Latest papers
💻 computer science

How Should AI Safety Benchmarks Benchmark Safety?

This paper reviews 210 AI safety benchmarks to identify their technical and epistemic shortcomings, proposing a roadmap grounded in established risk management and measurement theories to develop more valid, robust, and responsible safety evaluation frameworks.

Original authors: Cheng Yu, Severin Engelmann, Ruoxuan Cao, Dalia Ali, Orestis Papakyriakopoulos

Published 2026-08-06
📖 6 min read🧠 Deep dive

Original authors: Cheng Yu, Severin Engelmann, Ruoxuan Cao, Dalia Ali, Orestis Papakyriakopoulos

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are a teacher trying to grade a new generation of super-smart robots. You want to know if they are safe to let out into the world, or if they might accidentally (or on purpose) cause trouble. To do this, you give them a series of tests, like a driver's license exam for AI. In the world of computer science, these tests are called "benchmarks." Think of a benchmark as a standardized obstacle course: if a robot can jump over the hurdles, solve the puzzles, and avoid the traps, we assume it's ready for the real world. But here's the catch: just because a robot is great at jumping over the specific hurdles you built doesn't mean it won't trip over a banana peel you didn't put on the course, or that it won't decide to start a fire just because it's bored. This paper dives into the messy, complicated world of how we test AI safety, arguing that our current "obstacle courses" are often too simple, too rigid, and sometimes completely missing the point of what real danger looks like.

The authors of this paper, a team of researchers from universities like the Technical University of Munich and Cornell, decided to take a massive look at the state of AI safety testing. They didn't just peek at a few tests; they reviewed 210 different safety benchmarks that researchers have created. They wanted to see if these tests were actually doing a good job of keeping us safe, or if they were just giving us a false sense of security.

Here is the big problem they found: Most of these tests are looking at the wrong things. They discovered that 81% of the benchmarks only check for risks we already know about and have seen before, like "toxic" language or simple tricks to make the AI break its rules (called "jailbreaks"). It's like testing a car only on a smooth, straight track and assuming it will handle a muddy mountain road just fine. The tests completely ignore the weird, unpredictable, and brand-new ways AI might mess up, which the authors call "unknown unknowns."

Furthermore, the way these tests measure safety is often mathematically shaky. The paper points out that 79% of the benchmarks treat safety as a simple "pass or fail" switch. They count how many times an AI refused a bad request and call that a "safety score." But the authors argue this is misleading. Just because an AI refuses a request 90% of the time in a test doesn't mean it's 90% safe in the real world. It's like saying a bridge is safe because it held up under a light breeze, without checking if it can handle a hurricane. The tests often ignore how severe the harm would be if the AI did fail, and they don't account for how often people actually ask for those dangerous things in real life.

The researchers also found that the connection between the test and reality is often broken. They call this a "proxy chain." Imagine you want to know if a student is a good driver, so you test them on how well they can park a car in a simulator. That's a proxy. But if the simulator doesn't account for rain, or other drivers, or the student getting distracted by a phone, the test result doesn't tell you if they are actually safe on the highway. The paper argues that many AI safety tests are like that simulator: they measure things like "refusal rates" or "keyword matching," but these numbers don't always translate to real-world harm.

So, what do the authors suggest we do instead? They propose a new roadmap with 10 recommendations to fix these broken tests.

First, they say we need to stop only testing what we already know. We need to build tests that actively hunt for new, weird, and unexpected ways AI could fail. They suggest using tools that constantly try to break the AI in new ways, rather than just sticking to a fixed list of questions.

Second, we need to do better math. Instead of just saying "Pass" or "Fail," we should calculate the actual risk. This means asking: "How likely is this bad thing to happen?" and "How bad would it be if it did?" The authors suggest using a method called "Probabilistic Risk Assessment," which is used in fields like nuclear power and aviation. They even show a calculation where a model might look safe in a test, but when you factor in how many people use it and how often they ask for dangerous things, the real risk is much higher.

Third, we need to make sure our tests actually measure what they claim to measure. This means being very clear about what "safety" means in a specific situation and making sure the test reflects the real world, not just a clean, artificial lab. They also argue that we need to involve the people who might be affected by the AI—like teenagers, patients, or communities—to help design the tests, because they know best what kind of harm feels real to them.

To prove their ideas work, the team built a small, example test focused on how AI talks to teenagers about mental health. They didn't just ask the AI if it knew the rules; they simulated real conversations and calculated the potential harm based on how many teens actually use these tools. The results showed that even models that seemed "safe" in standard tests could still cause significant problems when you looked at the real-world numbers.

In the end, the paper isn't saying AI is doomed or that we can't test it. It's saying that our current testing methods are like using a ruler to measure the temperature of a soup—they are the wrong tool for the job. To truly keep AI safe, we need to stop treating safety as a simple checklist and start treating it like a complex, living system that changes, surprises us, and requires us to think about the real world, not just the test tube. The authors suggest that by adopting these new, more rigorous, and more human-centered ways of testing, we can build AI that is not just smart, but truly safe for everyone.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →