MentalBench: A DSM-Grounded Benchmark for Evaluating Psychiatric Diagnostic Capability of Large Language Models
This paper introduces MentalBench, a DSM-5-grounded benchmark utilizing a psychiatrist-validated knowledge graph to evaluate large language models' diagnostic capabilities, revealing that while models excel at knowledge retrieval, they struggle with confidence calibration in complex, ambiguous clinical scenarios.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
The Big Idea: A "Driver's License" Test for AI Doctors
Imagine you want to know if a self-driving car is ready for the road. You wouldn't just ask it, "Do you know what a stop sign looks like?" You would put it in a complex traffic jam, with fog, confusing signs, and pedestrians stepping out unexpectedly, to see if it can actually drive safely.
This paper does the exact same thing for Large Language Models (LLMs) trying to act as psychiatric assistants. The authors, a team of researchers and psychiatrists, built a new test called MENTALBENCH. Their goal wasn't to see if AI knows the definitions of mental illnesses, but to see if it can actually diagnose them in messy, real-world scenarios.
The Problem: AI is Good at Textbooks, Bad at Reality
Currently, most tests for AI in mental health are like asking a student to recite a textbook chapter.
- The Old Way: The AI is given a clean, perfect summary of a patient's symptoms (e.g., "Patient has sadness, sleep issues, and low energy for 3 weeks"). The AI says, "That's Depression."
- The Reality: Real patients don't speak in textbooks. They say things like, "I just feel empty, I can't get out of bed, and my boss is driving me crazy, but I don't know why." They might forget to mention a key symptom, or they might describe a symptom in a way that sounds like two different diseases.
The paper argues that previous AI tests were too easy because they didn't force the AI to handle this "messiness" or the tricky rules doctors use to tell similar diseases apart (like telling the difference between Bipolar Disorder and Depression).
The Solution: The "Mental Knowledge Graph" (MENTALKG)
To build a fair test, the researchers couldn't just guess what questions to ask. They needed a master blueprint.
- The Analogy: Imagine the DSM-5 (the big medical book doctors use) as a giant, dense library of rules written in complex legal language. It's hard for a computer to navigate.
- The Innovation: The team hired expert psychiatrists to turn that library into a MENTALKG (a Knowledge Graph). Think of this as a giant, interactive flowchart or a "choose your own adventure" map.
- It connects 23 different mental disorders.
- It maps out exactly which symptoms belong to which disease.
- Crucially, it maps out the rules for differentiation: "If the patient has symptom X, it's Disease A. But if they also have symptom Y for more than two weeks, it's actually Disease B."
This graph acts as the "truth" or the "answer key" for the entire experiment.
The Test: 24,750 Scenarios
Using this map, the researchers generated 24,750 synthetic patient cases. They didn't just make up random stories; they used the graph to ensure every story was medically accurate but varied in difficulty. They created four types of challenges:
- The "Perfect Report" (Type 1): A clean, professional medical summary. (Easy for AI).
- The "Confused Patient" (Type 2): A messy, first-person story where the patient forgets details or speaks in slang. (Hard for AI).
- The "Gray Area" (Type 3): A case where the symptoms fit two different diseases equally well because a key piece of evidence is missing. The correct answer is "It could be either."
- The "Tie-Breaker" (Type 4): A case where the symptoms fit two diseases, but one tiny detail (like the duration of a symptom) proves it is only one specific disease.
The Results: AI is Overconfident and Confused
When they ran the test on the smartest AI models available (including GPT-4o, Claude, and open-source models), the results were sobering:
- The "Textbook" Success: When the AI was given the clean, professional summaries (Type 1), it did great. It knew the definitions.
- The "Real World" Crash: When the AI had to read a messy, incomplete patient story (Type 2), its performance dropped significantly. It struggled to translate human language into medical rules.
- The "Differential Diagnosis" Failure: This was the biggest issue. When two diseases looked similar (Type 3 and 4):
- Open-source models tended to be over-diagnostic. They would say, "It could be A, B, and C!" even when the rules clearly ruled out B and C. They couldn't say "No."
- Closed-source (proprietary) models tended to be under-diagnostic. In ambiguous situations, they would force a single answer, saying "It's definitely A," even when the evidence was too weak to be sure. They couldn't say "I don't know."
The Conclusion
The paper concludes that while AI has memorized the vocabulary of psychiatry, it hasn't learned the logic of diagnosis. It struggles to handle the uncertainty and ambiguity that defines real human mental health.
Crucially, the paper states:
- This benchmark is not a tool for actual clinical diagnosis.
- The data is synthetic (made up by AI based on expert rules), not real patient records.
- The findings suggest that current AI is not reliable enough to be used as a decision-support tool in a real doctor's office, especially when the case is complex or the patient's story is incomplete.
In short: The AI knows the rules of the game, but it keeps losing when the players start breaking the rules or speaking in riddles. Before we trust AI with mental health, it needs to learn how to think like a doctor, not just a dictionary.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.