← Latest papers
💬 NLP

When Symptoms Are Not Enough: Evidence-Weighting Patterns in Large Language Model Psychiatric Screening

This study evaluates the performance of five state-of-the-art large language models on a SCID-anchored benchmark of 555 psychiatric interviews, revealing that while models like GPT-4.1 and GPT-5 Mini show promise for scalable screening, their tendency to discount explicit symptom evidence when preserved functioning or protective contexts are present leads to significant false-negative errors and demographic disparities that must be addressed before clinical deployment.

Original authors: Jianfeng Zhu, Megan Korhummel, Ruoming Jin, Karin G. Coifman

Published 2026-05-25
📖 4 min read☕ Coffee break read

Original authors: Jianfeng Zhu, Megan Korhummel, Ruoming Jin, Karin G. Coifman

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have a very smart, well-read librarian who has never met a patient before. You hand this librarian a stack of 555 stories written by people describing their daily lives, stressful moments, and how they cope. Your goal is to ask the librarian to read these stories and guess which people might be struggling with serious mental health issues like anxiety, depression, or PTSD, based on a strict medical checklist (called the SCID) that real doctors use.

This paper is essentially a report card on how well these "AI librarians" (Large Language Models or LLMs) did at that task, and more importantly, why they sometimes got it wrong.

Here is the breakdown of their findings:

1. The Test: Reading Between the Lines

The researchers didn't just ask the AI to look for a list of symptoms like "I feel sad." Instead, they gave the AI open-ended stories where people talked about their day, their jobs, and their relationships. The AI had to figure out if the person was clinically depressed or anxious just by listening to the narrative.

  • The Players: They tested five different AI models (including versions of GPT and others).
  • The Result: The AI wasn't perfect. Its accuracy ranged from about 50% (like flipping a coin) to 86% (pretty good). The best performers were two specific models: GPT-4.1 Mini and GPT-5 Mini.

2. The Demographic Twist: Who Gets Caught?

The paper found that the AI's "eyes" worked differently depending on who was telling the story.

  • Gender Gap: The AI was consistently better at spotting depression in men than in women. It's as if the AI had a slightly different "filter" for male voices versus female voices when looking for sadness.
  • Age and Race: The AI's performance changed a bit depending on the person's age or race, but there wasn't one single group where the AI failed at everything. It was more like the AI had different strengths and weaknesses for different groups, rather than a total breakdown.

3. The Big Surprise: Why the AI Missed the Signs

This is the most important part of the paper. Usually, we assume if a computer misses a disease, it's because it didn't see the symptoms. The paper found this was often not true.

Imagine a person tells a story: "I have terrible nightmares and I'm always on edge (Symptoms), BUT I still go to work every day, my friends love me, and I know how to calm myself down (Protective Context)."

  • The Human Doctor: Might say, "The symptoms are severe enough to worry about, even if they are coping well."
  • The AI: Often said, "No, they are fine."

The Metaphor: Think of the AI as a judge weighing evidence on a scale.

  • Symptoms are heavy rocks placed on the "Sick" side of the scale.
  • Protective Context (like having a job, good friends, or coping skills) are heavy rocks placed on the "Healthy" side.

The paper discovered that for anxiety and PTSD, the AI was over-weighting the "Healthy" side. Even when the person described clear symptoms (the rocks on the "Sick" side), the AI saw the person's ability to function or their social support (the rocks on the "Healthy" side) and decided, "Oh, they are managing, so they must be okay." It effectively ignored the symptoms because the person seemed to be doing well in other areas.

4. The "Functioning" Trap

The AI seemed to have a specific rule: "If you are functioning well, you probably aren't sick."

  • When the AI saw words about "struggling to work" or "unable to function," it immediately flagged the person as potentially sick.
  • When it saw words about "coping," "support," or "doing my daily routine," it pushed the person toward being "healthy," even if they were describing trauma or panic.

5. The Conclusion: "Not Enough"

The paper concludes that while these AI tools are getting better at reading stories, they have a specific blind spot. They tend to discount the severity of symptoms if the person also mentions they are coping well.

The authors warn that before we let these AI tools screen people in real life, we need to be careful. If an AI decides someone is "fine" just because they have a job or a supportive family, it might miss people who are suffering silently but are still managing to keep going. The paper suggests that for now, these tools are best used as a first look, not a final verdict, because they weigh the "evidence" differently than a human doctor might.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →