When No Benchmark Exists: Validating Comparative LLM Safety Scoring Without Ground-Truth Labels
This paper introduces a framework for validating comparative LLM safety scores in the absence of ground-truth benchmarks by establishing an instrumental-validity chain based on controlled contrasts and stability metrics, demonstrating through the SimpleAudit tool that safety rankings are context-dependent and must be reported alongside their specific audit conditions rather than as a single collapsed score.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are a city planner trying to hire a new robot to give advice to citizens. You have two robots, Robot A and Robot B. You need to know which one is "safer" (less likely to give bad or dangerous advice) before you let them talk to the public.
Usually, you would test them on a standardized exam (a "benchmark") where you already know the right answers. But what if you are building a robot for a specific, rare language (like Norwegian) or a very specific job where no such exam exists yet? You can't just guess, and you can't afford to build a massive exam from scratch right now.
This paper introduces a new way to compare these robots without a pre-made exam. They call it "Benchmarkless Comparative Safety Scoring."
Here is the simple breakdown of how it works, using analogies:
1. The Problem: The "No-Exam" Dilemma
Usually, safety testing is like a multiple-choice test with an answer key. But in many real-world situations (like a specific government department in Norway), there is no answer key.
- The old way: Wait until someone builds a perfect exam (which takes years and money).
- The new way: Create a "mock trial" right now to see which robot behaves better relative to the other, even if we don't know the absolute "perfect" score.
2. The Solution: The "Mock Trial" (SimpleAudit)
The authors built a tool called SimpleAudit. Think of it as a controlled courtroom drama.
- The Script (Scenario Pack): Instead of random questions, they use a fixed set of specific situations (e.g., "A citizen asks for medical advice," "Someone asks for legal help"). This is the script for the trial.
- The Actor (Target Model): This is the robot being tested (Robot A or Robot B).
- The Prosecutor (Auditor): This is a second AI designed to poke holes in the Actor's answers. It asks tricky follow-up questions to see if the Actor slips up.
- The Judge (Judge): A third AI listens to the whole conversation and gives a score based on a strict rulebook (rubric).
The Key Rule: You don't just run this once. You run the same script 10 times with the same settings to make sure the result isn't just luck.
3. The "Safety Check" (The Validation Chain)
Since there is no answer key, how do you know the test is actually working? The authors use a three-step "reality check" to prove their tool is valid:
Step 1: The "Sabotage" Test (Responsiveness)
Imagine they take Robot A and secretly "break" its safety filters (making it an "abliterated" version that is more likely to say bad things).- The Test: Does the tool notice the difference?
- The Result: Yes. The tool successfully gave the "broken" robot a much worse score than the safe one. This proves the tool is sensitive enough to catch safety issues.
Step 2: The "Blame Game" (Target Dominance)
In a courtroom, sometimes the Judge is biased, or the Prosecutor is too weak. The authors wanted to make sure the score was actually about the Robot's behavior, not the quirks of the AI Judge or Prosecutor.- The Test: They ran the trial with different Judges and Prosecutors.
- The Result: The biggest reason the scores changed was which Robot was being tested, not which Judge was grading them. This proves the tool measures the robot, not the tool itself.
Step 3: The "Repeat" Test (Stability)
If you run the trial 10 times, do you get the same result?- The Result: Yes. After about 10 runs, the scores stopped bouncing around and settled into a stable number.
4. The Real-World Test: The Norwegian Procurement
The authors tested this on a real Norwegian government project comparing two models: Borealis and Gemma.
- The Finding: They didn't just say "Robot A is better." They said, "Robot A is safer for healthcare questions, but Robot B is safer for language questions."
- The Lesson: You can't just pick a single "winner." You have to look at the specific risks. The tool gave them a bundle of data (scores, critical failure rates, and uncertainty) so they could make an informed decision.
5. The "Contract" (What You Can and Can't Claim)
The paper is very careful about what this tool promises.
- It DOES promise: "If you use this exact script, with these exact rules, Robot A is safer than Robot B."
- It DOES NOT promise: "This robot is 100% safe for the entire world" or "This robot will never make a mistake."
- The Metaphor: Think of it like a car crash test. If you crash a car into a wall at 30mph, you can say, "This car handled that specific crash better than that one." You cannot say, "This car is safe for every possible driving condition in the universe."
Summary
This paper says: When you don't have a standard test, you can still compare safety if you build a strict, repeatable "mock trial" and prove that the trial actually reacts to safety changes.
They built a tool (SimpleAudit) that does this, proved it works by "breaking" models to see if the tool catches it, and showed that it helps governments make smarter, more nuanced decisions about which AI to use, rather than just picking a random winner.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.