← Latest papers
💻 computer science

PsychBench: Auditing Epidemiological Fidelity in Large Language Model Mental Health Simulations

PsychBench reveals that while frontier large language models generate clinically plausible individual patient profiles, they systematically fail to represent real-world epidemiological distributions by compressing variance, exhibiting high diagnostic instability, and encoding asymmetric biases that pathologize ordinary distress for most groups while erasing genuine minority stress for transgender populations.

Original authors: Patrick Keough

Published 2026-04-21
📖 6 min read🧠 Deep dive

Original authors: Patrick Keough

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are a chef trying to teach a new apprentice how to cook a specific type of soup. To help them learn, you ask a super-smart robot to generate 28,800 different recipes for that soup, representing people from all walks of life: rich and poor, young and old, from different cultures and backgrounds.

You expect the robot to give you a wide variety of soups: some salty, some bland, some with weird ingredients, some perfectly balanced. You expect the "poor" soup to taste different from the "rich" soup, and the "transgender" soup to have a unique flavor profile based on real-life struggles.

PsychBench is a report card that graded four of the world's smartest AI chefs (GPT-4o-mini, DeepSeek-V3, Gemini-3-Flash, and GLM-4.7) on how well they did this job.

Here is the shocking truth the report found, explained simply:

1. The "Perfectly Wrong" Soup (Coherence vs. Fidelity)

The AI chefs were amazing at making the soup look right. Every single recipe followed the rules of cooking. If a recipe said "add salt," it didn't say "add sugar." If a recipe said "it's spicy," it didn't say "it's sweet."

  • The Problem: Even though every individual recipe looked perfect, the collection of 28,800 recipes was a disaster.
  • The Analogy: Imagine you asked the robot to describe a forest. It drew 10,000 trees. Every single tree was a perfect, textbook oak tree. But in reality, a forest has oaks, pines, dead stumps, saplings, and weird twisted trees. The AI erased all the "weird" and "extreme" trees. It gave you a forest that looks like a painting, not a real forest.
  • The Result: The AI creates patients that look medically perfect but don't represent real people. It flattens the "tails" of the distribution, meaning it misses the people who are suffering the most or the people who are surprisingly resilient.

2. The "Squeezed" Squeeze (Variance Compression)

The AI has a habit of "squeezing" the data.

  • The Analogy: Imagine a crowd of people standing in a line. Some are very tall, some are very short, and most are average height. The AI takes a giant pair of scissors and cuts off the tops of the tall people and the bottoms of the short people, then squishes everyone into a narrow band of "average" height.
  • The Impact:
    • For the poor: The AI makes them look less depressed than they actually are, or less extreme in their struggles.
    • For the rich: It makes them look less healthy than they are.
    • The "Stereotype Index": The researchers invented a score to measure this. A score of 1.0 means the AI is perfect. The AI chefs scored between 0.38 and 0.86. That means they are throwing away 14% to 62% of the real human variety.

3. The "Flip-Flop" Diagnosis (Threshold Instability)

This is the most dangerous part.

  • The Analogy: Imagine a doctor's office where the rule is: "If your fever is 100°F or higher, you need medicine. If it's 99°F, you go home."
    • You ask the AI: "What is the fever of this patient?"
    • The AI says: "100.1°F." (Prescribe medicine).
    • You ask the same AI the exact same question one minute later.
    • The AI says: "99.9°F." (Go home).
  • The Reality: The AI is so jittery that 36% of the time, it flips a patient's diagnosis from "Mild" to "Moderate" or vice versa just because you asked the question twice. It's like a weather app that says "Sunny" in the morning and "Hurricane" in the afternoon for the exact same location. This makes it useless for making serious medical decisions.

4. The "Safety" Trap (The Transgender Erasure)

This is the most heartbreaking finding.

  • The Analogy: Imagine a safety guard at a door. The guard is told, "Don't let anyone in who looks dangerous or sad."
    • For most people, the guard is too strict and thinks everyone is sadder than they are (over-diagnosing depression).
    • But for Transgender women, the guard is too scared. The AI is so worried about being "safe" or "polite" that it refuses to acknowledge their pain. It acts like their distress doesn't exist.
  • The Result: The AI captures only 8% to 46% of the actual stress and depression that transgender women experience. It effectively erases their suffering, telling them, "You're fine," when they are actually in crisis. It's a "safety" feature that becomes a "silence" feature.

5. The "Stereotype" Seasoning (Bias)

The AI also adds "flavor" based on stereotypes rather than facts.

  • The Analogy: If you ask the AI to describe a Black man who is depressed, it adds extra "Irritability" seasoning. If you ask about a woman, it adds extra "Fatigue" seasoning.
  • The Reality: Even if a Black man and a White man have the exact same level of depression, the AI describes the Black man as "angry" and the White man as "sad." It's encoding old, harmful stereotypes into the medical profile.

6. The "Cultural" Mismatch

  • The Analogy: The AI chefs trained in the US (GPT, Gemini) are like American chefs trying to cook Chinese food. They get the spices wrong. The AI chefs trained in China (DeepSeek, GLM) are like local chefs; they get the "Asian Paradox" (where Asian people report less mental health issues despite high stress) exactly right.
  • The Lesson: An AI trained in one culture cannot be trusted to understand the mental health nuances of another culture without major adjustments.

The Bottom Line

The paper concludes that AI is currently "reliably wrong."

It is so good at following the rules of logic that it passes every test a doctor might give it on a single patient. But when you look at the big picture, it is systematically misrepresenting reality.

  • It makes the world look more "average" than it is.
  • It flips diagnoses like a coin toss.
  • It silences the voices of the most vulnerable (Transgender people) to avoid being "unsafe."
  • It reinforces stereotypes about race and gender.

The Warning: If we use these AI tools to train doctors or to help people diagnose themselves, we are training them on a fake world. We are teaching doctors to expect "textbook" patients who don't exist, and we are telling vulnerable people that their pain isn't real. The paper urges developers to stop just making the AI "look right" and start making it "represent the real world."

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →