← Latest papers
💬 NLP

Ask Before You Diagnose: Safe-Psych, a Sequential Evaluation Benchmark for LLMs in Psychiatry

The paper introduces Safe-Psych, a sequential benchmark using over 1,000 real-world psychiatric notes to reveal that even state-of-the-art large language models struggle to recognize incomplete clinical information, frequently making premature diagnoses or failing to request clarification despite safety-aware prompting.

Original authors: Oriana Presacan, Andreea Grama, Larisa Irimină, Alireza Nik, Jaya Ojha, Vajira Thambawita, Ciprian I. Băcilă, Bogdan Ionescu, Michael A. Riegler

Published 2026-07-16
📖 4 min read☕ Coffee break read

Original authors: Oriana Presacan, Andreea Grama, Larisa Irimină, Alireza Nik, Jaya Ojha, Vajira Thambawita, Ciprian I. Băcilă, Bogdan Ionescu, Michael A. Riegler

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are a detective trying to solve a mystery. In the world of artificial intelligence, there are super-smart computer programs called Large Language Models (LLMs). Think of them as incredibly well-read detectives who have read almost every book, article, and note ever written. They are great at answering questions when they have all the facts right in front of them. But in the real world, especially in places like hospitals, detectives rarely get the whole story at once. They get a clue here, a witness statement there, and maybe a piece of evidence that doesn't quite fit yet.

The big question scientists are asking is: Can these AI detectives know when they don't have enough clues to solve the case? A truly smart detective shouldn't just guess the culprit because they are confident; they should say, "Wait, I need to ask one more question," or "I can't solve this yet." This is called "calibration"—knowing what you know and, more importantly, knowing what you don't know. If an AI guesses a medical diagnosis too early, it could lead to serious mistakes. But if it refuses to answer even when it could have solved it, it becomes useless. Finding the perfect balance is the holy grail of making AI safe for healthcare.

Enter Safe-Psych, a new experiment designed to test exactly this skill in the tricky world of psychiatry. The researchers created a special game where they didn't give the AI the whole patient file at once. Instead, they acted like a real doctor, revealing the patient's story piece by piece: first just the symptoms, then the history, then the exam results, and so on. At every step, the AI had to decide: "Do I have enough to make a diagnosis?" "Should I ask for more info?" or "Should I admit I'm stuck?"

The results were a bit of a wake-up call. The paper found that even the most advanced AI models, the "super-detectives" of the tech world, are terrible at knowing when to stop and ask for help. When the information was incomplete, these models kept guessing anyway. In fact, for most of the models tested, they tried to diagnose cases where they didn't have enough evidence more than 60% of the time. It's like a detective shouting, "It was the butler!" after only hearing the front door open, without ever checking the living room.

The researchers also tried to teach the AI to be more careful by giving it a special instruction: "Only guess if you are sure, otherwise say 'I don't know'." This helped a little, but it created a new problem. The AI swung too far the other way. Instead of guessing too much, it started refusing to guess even when it could have been right. It's as if the detective, afraid of making a mistake, decided to never solve any cases at all. The study suggests that simply telling an AI to "be safe" doesn't actually teach it how to recognize when information is missing; it just shifts the errors from "guessing wrong" to "refusing to guess."

Perhaps the most interesting finding was about timing. When the AI waited until it had seen all the clues before making a call, its answers were much better. But when it rushed to diagnose after seeing just the first few sentences, its accuracy dropped significantly. The paper shows that capability doesn't equal safety: just because a model is smart enough to get the right answer when it has the full story, doesn't mean it knows when it has the full story.

In short, Safe-Psych reveals that our current AI detectives are still too eager to solve the mystery. They need to learn the hardest lesson of all: sometimes, the most important thing a detective can do is admit they need more clues before they point a finger. Until they learn that, we can't fully trust them to help doctors make life-or-death decisions.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →