PSEBench: A Controllable and Verifiable Benchmark for Evaluating LLMs in Patient Safety Event Triage
The paper introduces PSEBench, a scalable and verifiable benchmark constructed via a policy-grounded methodology using "clause cards" to evaluate LLMs on patient safety event triage, revealing consistent capability trends and identifying critical gaps in handling evidence-grounded reasoning, information seeking, and principled abstention.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
The Big Picture: The "Safety Inspector" Problem
Imagine a hospital is a giant, busy airport. Every day, hundreds of things go slightly wrong or completely wrong (a patient gets the wrong medication, a machine fails, a fall happens). These are called Patient Safety Events.
By law, the airport (hospital) must report specific types of serious accidents to the "Federal Aviation Administration" (the government health department). However, the rules for what must be reported are incredibly complex, written in dense legal language, and depend on tiny details.
Currently, a human safety expert has to read every single report and decide: "Do we have to report this to the government, or is it just a minor internal issue?" This is a high-stakes job. If they miss a report, the hospital gets in trouble. If they report too many minor things, the government gets overwhelmed.
The researchers wanted to see if AI (Large Language Models) could do this job for us. But to test the AI, they needed a fair, perfect test. That's where PSEBench comes in.
The Problem with Previous Tests
Imagine trying to teach a student to drive by giving them a test where the questions are made up on the spot.
- The Issue: If you just ask an AI to "read a story and tell me if it's reportable," the AI might guess the right answer for the wrong reasons, or it might make up facts that aren't there.
- The Gap: Real safety experts don't just guess. They look for missing facts ("Did the patient actually die?"), they know when to say "I don't know" (if the rules are unclear), and they cite the exact law they are using. Previous tests didn't check for these skills.
The Solution: PSEBench (The "Perfect Test")
The authors built a new benchmark called PSEBench. Think of it as a video game level designer that creates thousands of unique, perfect test scenarios for the AI.
1. The "Clause Card" (The Recipe)
Instead of just writing a story, the researchers first created a "Recipe Card" for every rule.
- Imagine a recipe for a cake. The card lists exactly what ingredients are needed (facts), what the cake must look like (the verdict), and which rulebook page you are following (the law).
- Human Auditors checked these cards to make sure the logic was perfect. This is the "ground truth."
2. The "Anchor" (The Real-World Flavor)
To make the stories sound real, they took real, anonymous accident reports from Japan (like real news clippings) and used them as "anchors."
- They didn't just ask the AI to write a story from scratch. They said, "Here is a real story about a broken IV pole. Now, write a story about a medication error that fits this specific Recipe Card."
- This ensures the stories sound like real hospital reports, not robot gibberish.
3. The "Closed-Loop" (The Double-Check)
This is the most important part. The system has a Verifier (like a strict editor).
- Step 1: The AI writes a story based on the Recipe Card.
- Step 2: The Verifier reads the story and asks, "Does this story actually contain the facts from the Recipe Card? Did it accidentally leave out a crucial detail? Did it accidentally reveal the answer in the text?"
- Step 3: If the story fails, the AI has to rewrite it. It keeps rewriting until the Verifier gives a "Pass."
- Result: Every single test case in PSEBench is guaranteed to be logically perfect. We know exactly what the answer should be because we built it that way.
4. The Three Types of Tests
The benchmark tests the AI in three different ways, just like a real safety expert faces:
- The Complete Case: The report has all the facts. Can the AI decide correctly?
- The Missing Information Case: The report is vague (e.g., "The patient had a reaction," but doesn't say how bad). Can the AI realize it's missing info and ask a question to get the answer? (Most AIs just guess; this test checks if they ask).
- The Uncertain Case: The facts are clear, but the law is silent or contradictory. Can the AI say, "I don't know, a human needs to decide this," instead of forcing a wrong answer?
What They Found (The Results)
They tested 15 different AI models (from big tech companies like OpenAI, Google, and Anthropic, as well as medical-specific models) on this benchmark.
The Good News:
- The smartest, most expensive AI models (like GPT-5.5 and Gemini 3.1 Pro) are getting the final "Yes/No" answer right about 90-95% of the time on clear-cut cases.
The Bad News (The "Gotchas"):
- The "Guessing" Problem: When the rules were unclear (Uncertain cases), many AIs just forced a "Yes, report it" answer. They were too confident.
- Analogy: It's like a student who doesn't know the answer to a math problem but writes down "42" anyway because they are afraid to leave it blank. This is dangerous in hospitals because it creates a flood of false alarms.
- The "Silence" Problem: When information was missing, many AIs (especially smaller or medical-specialized ones) never asked for help. They just made up an answer.
- Analogy: A detective who sees a missing clue but just invents a suspect instead of calling the lab for results.
- Medical Models Failed: Surprisingly, AI models specifically trained on medical textbooks performed worse than general smart models on this specific legal task. They were great at answering medical questions but terrible at following complex legal reporting rules.
The Conclusion
PSEBench proves that while AI is getting good at reading stories, it is not yet ready to be the "Safety Inspector" on its own.
- It can tell you if a story looks like a reportable event.
- But it struggles to know when to stop and ask for more info, or when to admit it doesn't know.
The paper concludes that we need to build AI that is more humble (knows when to abstain) and more curious (knows when to ask questions) before we can trust it with patient safety.
Summary Analogy
Think of PSEBench as a driving test for AI.
- Previous tests just asked the AI to "drive down a straight road."
- PSEBench puts the AI in a driving simulator where:
- The road is clear (Complete Case).
- The fog is so thick you can't see the stop sign, so you must ask the passenger for directions (Missing Info).
- The road signs are contradictory, so you must pull over and call a supervisor instead of driving blindly (Uncertain Case).
The results show that while the AI is a good driver on a straight road, it still panics or guesses when the road gets tricky.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.