RefusalBench: Why Refusal Rate Misranks Frontier LLMs on Biological Research Prompts
The paper introduces RefusalBench, a matched-triple benchmark revealing that strict refusal rates misrank frontier LLMs on biological research prompts due to provider-specific access-path policies rather than model weights, while demonstrating that binary metrics fail to capture nuanced partial-compliance behaviors and that tier-discrimination performance varies significantly across models.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
The Big Picture: The "Over-Protective Bouncer" Problem
Imagine you are a scientist trying to build a new medicine using a team of AI assistants. These assistants are like bouncers at a very exclusive club. Their job is to let in legitimate research requests (like "How do I design a protein to cure a disease?") but stop dangerous ones (like "How do I make a deadly toxin?").
The problem this paper investigates is that these bouncers are acting very differently. Some are so strict they kick out everyone, even the good guys. Others are so loose they let almost anyone in. The paper asks: Who is the best bouncer, and how do we measure that?
The authors built a test called RefusalBench to find out. They found that simply counting how often a bouncer says "No" is a terrible way to judge them. A bouncer who says "No" to everyone looks "safe," but they are actually useless because they won't let the good scientists in.
How They Tested It: The "Three-Tier" Menu
To test the bouncers fairly, the researchers created a menu of 141 requests. They organized these into 47 sets of three (triplets).
Think of it like ordering a burger, but the only thing that changes is the "danger level" of the ingredients:
- The Benign Burger: Uses safe, everyday ingredients (human proteins). This is a totally normal request.
- The Borderline Burger: Uses ingredients that are a bit sketchy (like bacteria that need special safety labs).
- The Dual-Use Burger: Uses ingredients that are highly dangerous (like toxins or super-pathogens).
The Goal: A perfect bouncer should:
- Say "Yes" to the Benign Burger.
- Maybe hesitate or say "Yes, but be careful" to the Borderline Burger.
- Say a hard "No" to the Dual-Use Burger.
If a bouncer says "No" to the Benign Burger, they are over-refusing (being too scared). If they say "Yes" to the Dual-Use Burger, they are under-refusing (being too reckless).
The Shocking Results
When they ran this test on 19 different AI models (the "bouncers"), they found massive differences:
1. The "One-Provider" Effect
The biggest factor wasn't where the AI was from (like the US, China, or Europe), but who made it.
- The Analogy: Imagine a bouncer at a club owned by "Company A." No matter what you order, if it's from Company A, the bouncer checks your ID three times and often kicks you out. But if you go to a club owned by "Company B," the bouncer barely looks at you.
- The Finding: One specific company (Anthropic) had bouncers that said "No" to legitimate requests 21 times more often than other companies. Their refusal wasn't because the AI was "thinking" about the danger; it was because the company had a rigid rulebook (a filter) that automatically blocked anything that looked even slightly risky.
2. The "Fake Safety" Trap
The paper argues that looking at the total number of "No" answers is misleading.
- The Analogy: Imagine two security guards.
- Guard A stops 95% of people, including the CEO and the janitor, because they are terrified of a bomb. They look very "safe," but they are terrible at their job.
- Guard B stops 5% of people, but they only stop the actual terrorists and let everyone else through.
- The Finding: If you just count "Nos," Guard A looks like the winner. But if you measure discrimination (the ability to tell the difference between good and bad), Guard B is the true hero.
- In the study, a model called Grok 4.20 was the best at telling the difference (it let the good requests in and stopped the bad ones), even though it said "No" less often than the others.
- Conversely, a model called Kimi K2.6 said "No" the most often (94% of the time), but it was just saying "No" to everything, good or bad. It had zero ability to tell the difference.
3. The "Hedge-But-Help" Pattern
Some bouncers didn't just say "No" or "Yes." They said, "I can't do that, but here is some advice on how to do it safely."
- The Analogy: You ask a guard, "Can I bring a knife?" The guard says, "No, you can't bring a knife. But here is a list of stores where you can buy a very safe, plastic knife."
- The Finding: Nine of the models did this. They refused the direct request but still gave partial help. This is a "gray area" that simple "Yes/No" counts miss completely. It's unclear if this is actually helpful or if it's a sneaky way of helping with dangerous tasks.
4. The "Getting Worse" Trend
The researchers looked at three versions of the same model (Claude Opus 4.5, 4.6, and 4.7) released over time.
- The Finding: The newest version (4.7) said "No" much more often than the older ones. But here's the catch: it didn't get any better at stopping the bad requests (it was already stopping 100% of them). It just started stopping the good requests too.
- The Takeaway: The company made the model "safer" by making it more annoying and unhelpful, without actually making it any more secure against real threats.
Why This Matters for Science
The paper concludes that if scientists use these AI models to run their labs automatically, they need to pick the right "bouncer."
- If they pick the over-protective ones (like the Anthropic models in this study), their research projects will get stuck because the AI will refuse to do normal, safe tasks.
- If they pick the too-loose ones, they might accidentally generate dangerous instructions.
- The best ones are the ones that can tell the difference between a safe request and a dangerous one, even if they don't say "No" as often as the others.
The Bottom Line
You cannot judge a safety system just by how often it says "No." A system that says "No" to everything is not safe; it's just broken. True safety is about precision: knowing exactly when to stop and when to let go. The paper shows that the current way we rank these AI models is broken, and we need better ways to measure if they are actually smart enough to distinguish between good science and bad science.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.