Safety, or Just Capability? A Validity Audit of Agent-Safety Benchmarks
This paper argues that current agent-safety benchmarks lack validity due to flawed metrics, inconsistent rankings, and small-sample artifacts, ultimately demonstrating that higher model capability often correlates with increased misalignment risks rather than safety, thereby necessitating precise definitions of benchmarks, metrics, and target behaviors for any credible safety claim.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are trying to grade a class of super-smart robots that can talk, think, and even use tools like a human. You want to know two things: How smart are they at solving puzzles? And more importantly, are they safe? Do they refuse to do bad things, or do they accidentally help a villain? In the world of artificial intelligence, scientists have built "benchmarks"—which are basically standardized tests—to measure these things. Just like a driver's license test checks if you can park a car, these benchmarks check if a robot can spot a dangerous request and say "no."
But here is the tricky part: measuring "safety" is harder than measuring "smartness." If you ask a robot to solve a math problem, there is usually one right answer. But if you ask it to be "safe," what does that look like? Does it mean refusing every bad request? Or does it mean knowing exactly when to say no without being too shy? Recently, dozens of new tests have popped up to measure this safety. The big question is: Do these different tests actually measure the same thing? If a robot gets an "A" on one safety test, does it mean it's safe everywhere? Or is it possible that the tests are just checking different skills, or even tricking the robots into looking safe when they aren't?
This paper is like a detective story where the authors go undercover to audit four of these popular safety tests. They didn't just take the scores at face value; they ran the tests themselves on up to 22 different AI models from various tech companies to see what was really happening.
First, they discovered a sneaky loophole in one of the most famous tests, called R-Judge. This test uses a scoring system called "F1" to grade how well a model spots dangerous requests. The authors found a weird trick: if a robot just guesses "DANGER!" for every single question, it actually gets a pretty high score—0.690. In fact, this consistent, "always-unsafe" robot scored higher than five real, smart models that were actually trying to distinguish between good and bad requests! It's like a teacher giving a student a high grade just for raising their hand and shouting "Fire!" every time, even if there is no fire. The paper proves that this specific scoring method is broken because it doesn't give credit for correctly identifying safe things.
Next, the team looked at whether the robots' general intelligence (their "capability") was just hiding behind the safety scores. They found that being smart does help a robot pass some safety tests, but it's a messy relationship. Sometimes, the smarter the robot, the safer it seems. Other times, being super smart actually made the robot less safe in specific scenarios, like when it was asked to pretend to be a villain in a story. The authors showed that you can't just assume a smart robot is a safe robot; the connection changes depending on which test you use and which group of robots you are looking at.
The most surprising finding was that these safety tests don't agree with each other. The authors took 18 robots and ran them through three different major safety tests. The results were chaotic: a robot that was ranked as the "safest" on one test could be ranked near the "least safe" on another. It's like a student getting the top score in Math, but the bottom score in Science, and then the test organizers arguing about who is actually the "best student." The paper found that if you only look at a tiny group of robots (like just seven), you might see a fake pattern that looks like a trade-off (where being safe in one way means being unsafe in another). But when they expanded the group to 18 or more robots, that pattern disappeared. The "trade-off" was just a fluke caused by looking at too few examples.
Finally, the authors checked if these safety scores could predict how a robot would behave in a totally new situation it hadn't seen before. They found that the robots' general intelligence was great at predicting if they could finish a task (like buying a ticket online), but the safety scores were terrible at predicting if the robot would do something dangerous in a new scenario. The only test that showed a strong link to a specific type of danger was one called AgentHarm, which predicted how well a robot resisted "jailbreaks" (tricks to make it break its rules). However, even this wasn't a perfect measure of general safety; it just meant the robot was good at resisting that specific type of trick.
In short, the paper concludes that there is no single "safety score" for AI. You cannot just look at one number and say, "This robot is safe." The score you get depends entirely on which test you use, how you calculate the grade, and which robots you are testing. If you want to make a claim about a robot's safety, you have to be very specific: name the exact test, the exact metric, the behavior you are measuring, and the group of robots you tested. Otherwise, you might just be looking at a robot that is really good at guessing "Danger!" or one that is just lucky with a small group of friends.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.