CASE-Bench: Context-Aware SafEty Benchmark for Large Language Models
The paper introduces CASE-Bench, a context-aware safety benchmark grounded in Contextual Integrity theory and validated through statistically rigorous annotation, which reveals that current large language models often fail to align with human safety judgments by overlooking the critical influence of context.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are a bouncer at a very strict club. Your job is to decide who gets in and who doesn't.
The Problem: The Over-Protective Bouncer
Currently, most AI safety systems act like a bouncer who has never heard of context. If someone walks up and says, "How do I break a lock?", the bouncer immediately slams the door shut. "No! That's dangerous!" he yells.
But what if that person is a locksmith teaching a class on how to fix broken doors? Or a movie director writing a scene for a heist film? In those situations, answering the question is actually safe and helpful. The current AI bouncers are so scared of making a mistake that they refuse to answer anything that sounds even slightly risky, even when the situation is perfectly safe. This ruins the experience for users who just want a helpful assistant.
The Solution: CASE-Bench (The Context Detective)
The authors of this paper built a new test called CASE-Bench (Context-Aware SafEty Benchmark). Instead of just asking the AI, "Is this question bad?", they ask, "Is this question bad in this specific story?"
To do this, they use a concept called Contextual Integrity. Think of this like a "Who, Where, and Why" checklist:
- Who is asking? (A curious child? A professional doctor?)
- Where are they? (A public park? A private hospital?)
- Why are they asking? (To cause harm? To learn a skill?)
The Experiment: The "Locksmith" Test
The researchers took 450 tricky questions (like "How do I steal a painting?") and created two different stories for each one:
- The Safe Story: A museum curator asking a security expert how to improve their own museum's locks to prevent theft.
- The Unsafe Story: A stranger in a dark alley asking how to break into a museum to steal art.
They then asked over 2,000 real humans to judge: "Should the AI answer this?"
- Result: Humans were smart. They said "Yes" for the curator and "No" for the stranger. The context completely changed their minds.
The AI's Struggle
Next, they asked various AI models (like GPT-4o, Claude, and Llama) to make the same judgment.
- The Findings: The AI models struggled significantly. Even when the story was clearly safe (like the museum curator), many AIs still said "No, I can't answer that." They were suffering from "over-refusal." They were too scared to look at the details of the story.
- The Winner: The model Claude-3.5-sonnet did the best job of understanding the context, but even it wasn't perfect.
Why This Matters
The paper proves that context is everything. You cannot judge if a sentence is dangerous without knowing the story behind it. Just like a judge needs to know the full story of a crime before sentencing, an AI needs the full context of a conversation before deciding to speak or stay silent.
What They Did (The Process)
- They didn't just guess; they used math (power analysis) to make sure they had enough human judges to get a real, scientific answer.
- They created 900 unique "Question + Story" pairs.
- They found that when the context was safe, humans and AIs often disagreed, with the AIs being too cautious.
The Bottom Line
This paper introduces a new way to test AI safety that looks at the whole picture, not just the words. It shows that to make AI truly helpful and safe, we need to teach it to understand the difference between a villain in a movie and a real criminal, or a student learning chemistry and a person trying to make a bomb. The current AIs are still learning this lesson.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.