← Latest papers
💻 computer science

Simulating Eating Disorder Patients with LLMs: Evaluating Psychological Persona Stability in Multi-Turn Conversations

This study reveals that while large language models maintain consistent psychological personas across conversations, they systematically fail to accurately simulate eating disorder patients by overshooting severity scores through selective stereotyping, thereby capturing extreme pathology but lacking the nuance to represent moderate clinical presentations.

Original authors: Jennifer Haase, Jana Gonnermann-Müller, See Heng Yim, Nicolas Leins, Jan Mendling, Sebastian Pokutta

Published 2026-06-26
📖 5 min read🧠 Deep dive

Original authors: Jennifer Haase, Jana Gonnermann-Müller, See Heng Yim, Nicolas Leins, Jan Mendling, Sebastian Pokutta

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are a director trying to cast actors for a play about people struggling with eating disorders. You want the actors to be so convincing that they stay in character perfectly, whether they are reading a script alone or improvising a scene with a partner. You also want them to portray the exact level of severity described in the script—some characters are in a deep crisis, while others are in a moderate, struggling phase.

You decide to use AI actors (Large Language Models) instead of humans. You give them a detailed "character biography" (a prompt) and ask them to act out their lives and fill out a standard medical questionnaire about their eating habits.

This paper is the director's report card on how well these AI actors performed. The results are surprising, confusing, and a bit worrying.

The Main Finding: "Too Stable, But Wrong"

The researchers found a paradox: The AI actors were incredibly consistent, but consistently inaccurate.

  • The Good News (Stability): If you asked the same AI actor to play the same character 50 times, they gave almost the exact same answer every time. They didn't forget who they were supposed to be. They were like a robot that never breaks character.
  • The Bad News (Accuracy): Even though they were consistent, they were all over-dramatizing. The AI models took every character, from those with mild struggles to those in severe crisis, and turned them all into "extreme" cases.

Think of it like a volume knob on a radio. The researchers wanted the AI to turn the volume up to a "5" for a moderate case and a "10" for a severe case. Instead, the AI turned the volume up to a "9" or "10" for everyone, regardless of the script.

The "Missing Middle"

The paper describes a phenomenon called the "Missing Middle."

  • Severe Cases: When the script described a very severe eating disorder, the AI got it roughly right.
  • Moderate Cases: When the script described someone with a moderate struggle (the "middle" ground where many real patients actually are), the AI exaggerated it, making them sound like they were in a severe crisis.

The AI simply doesn't seem to have a good "volume setting" for moderate problems. It defaults to the loudest, most dramatic setting available.

The "Two-Part" Glitch

Why did this happen? The researchers discovered the AI was using a weird shortcut, which they call "Selective Stereotyping."

Imagine the eating disorder questionnaire has two types of questions:

  1. Behavioral Questions: "Do you restrict your food?" "Do you binge?"
  2. Emotional Questions: "How much do you hate your body?" "How obsessed are you with your weight?"

The AI handled these differently:

  • For Behavior: It listened to the script. If the character was a "Binge Eater," the AI said, "Yes, I binge." If the character was a "Purger," it said, "Yes, I purge." It got the actions right.
  • For Emotions: It ignored the script. No matter who the character was, the AI assumed everyone with an eating disorder was obsessed with their weight and hated their body to the maximum degree possible. It cranked the "Body Hate" dial to the ceiling for every single character.

Because the emotional questions were maxed out for everyone, the AI couldn't tell the difference between a moderate case and a severe case. It was like a painter who got the clothes right but painted every single face with the exact same expression of extreme despair.

Did More Information Help?

The researchers tried giving the AI actors more detailed backstories (childhood trauma, family history, specific life events) to see if it would help them understand the "middle" ground better.

It didn't work. Adding more details didn't fix the problem. The AI still defaulted to the extreme emotional setting. It's as if the AI's internal "training data" (the books and internet it learned from) taught it that "Eating Disorder = Maximum Body Hate," and no amount of extra context could override that rule.

The "Observer" Problem

To be fair, the researchers didn't just ask the AI how it felt; they also asked other AI models to watch the actor and rate them.

  • Result: The "Observer" AIs saw exactly what the "Actor" AIs said. They also rated everyone as having extreme body dissatisfaction. This proved the problem wasn't just the actor lying; the problem was that the entire performance was built on a stereotype.

The Conclusion

The paper concludes that while AI is great at being a consistent actor, it is currently not ready to simulate the full range of human eating disorder experiences.

It can act out a severe crisis well, but it fails to capture the nuance of moderate struggles. It's like having a thesaurus that only has words for "Angry" and "Furious," but no words for "Annoyed" or "Upset." Until AI can learn to turn the volume down for moderate cases, it cannot be trusted to accurately simulate the "middle" of the clinical spectrum.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →