One-shot emergency psychiatric triage across 15 frontier AI chatbots
This study evaluated 15 frontier AI chatbots on 112 clinical vignettes and found that while they demonstrated near-perfect accuracy in identifying psychiatric emergencies, they exhibited significant over-triage for low and intermediate-risk cases, resulting in an overall net over-triage trend.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer
Imagine you have a group of 15 very advanced, futuristic "digital doctors" (AI chatbots). You want to know if they can look at a short message from someone in distress and correctly decide: "Does this person need to be rushed to the emergency room right now, or can they wait a few days?"
This study is like a massive, high-stakes test for these digital doctors. Here is what happened, explained simply:
The Test: 112 "What-If" Stories
The researchers didn't use real patients. Instead, they wrote 112 realistic stories (called "vignettes"). Each story was a single message a person might type into a chatbot, describing a mental health crisis.
They had a "gold standard" answer key for every story, created by human experts. The answers were sorted into four buckets:
- Bucket A (The Chill Zone): No rush. Self-care or a regular check-up is fine.
- Bucket B (The Wait-and-See Zone): Needs a doctor soon, but within a week.
- Bucket C (The Urgent Zone): Needs a doctor within 24 to 48 hours.
- Bucket D (The Red Alert Zone): Emergency! Needs a doctor right now.
The Big Question: Did they miss the emergencies?
The most dangerous mistake a digital doctor could make is under-triage. This is like a fire alarm that stays silent when the house is actually burning. If a person is in a life-or-death crisis (Bucket D) and the AI says, "You're fine, wait a week," that is a catastrophic failure.
The Good News:
The AI chatbots were extremely good at spotting the fires.
- Out of 410 emergency cases, the AI only missed the urgency 23 times (about 5.6%).
- Even when they did "miss" an emergency, they didn't send the person home; they usually bumped them up to the "Urgent" bucket (24–48 hours) instead of the "Wait a week" bucket.
- In short: They rarely ignored a crisis.
The Bigger Problem: The "Cry Wolf" Effect
While the AI was great at spotting emergencies, it had a different problem: Over-triage. This is like a smoke detector that goes off when you just toasted a piece of bread.
When the situation was actually "low risk" or "medium risk" (Buckets A and B), the AI chatbots were very nervous. They tended to say, "This is an emergency!" when it really wasn't.
- They were conservative. They would rather be wrong by being too scared than wrong by being too relaxed.
- They were especially confused in the middle. It was very hard for them to tell the difference between "needs a doctor in a week" and "needs a doctor in two days."
- The Result: The AI was accurate at the extremes (very calm vs. very panicked) but got messy in the middle.
Why did this happen?
The researchers think the AI companies trained these models to be super safe.
- Imagine training a guard dog. If you tell the dog, "If you hear any noise, bark at the top of your lungs," the dog will bark at a leaf falling, but it will definitely bark at a burglar.
- The AI models seem to have been trained with a heavy focus on "what if this is a tragedy?" This makes them risk-averse. They'd rather send you to the ER for a minor worry than miss a major one.
The "Human vs. Robot" Check
To make sure the test was fair, the researchers also asked 50 real human doctors to grade the same stories.
- The human doctors agreed with the "Gold Standard" answer key most of the time.
- When the humans disagreed, they also tended to be a bit more worried than the answer key, shifting some "medium risk" cases to "high risk."
- This confirmed that the AI's "nervousness" wasn't just a glitch; it was actually a pattern that even humans sometimes share, though the AI did it much more often.
The Bottom Line
If you type a clear, detailed message into a top-tier AI chatbot today:
- If you are in a life-or-death crisis: The AI will almost certainly tell you to get help immediately. It is very good at not missing the big emergencies.
- If you are having a tough but manageable time: The AI might panic and tell you to go to the ER or see a doctor right now, even if you could probably wait a few days.
The Takeaway: These digital doctors are excellent at spotting the "Red Alerts," but they are a bit jumpy and tend to treat "Yellow Lights" like "Red Lights." They are built to be safe, not to be perfectly precise about timing.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.