← Latest papers
💬 NLP

Evaluating the Presence of Sex Bias in Clinical Reasoning by Large Language Models

This study reveals that contemporary large language models exhibit significant, model-specific sex biases in clinical reasoning when processing non-informative patient vignettes, underscoring the critical need for conservative configuration, data auditing, and human oversight to ensure safe healthcare integration.

Original authors: Isabel Tsintsiper, Sheng Wong, Beth Albert, Shaun P Brennecke, Gabriel Davis Jones

Published 2026-02-05
📖 4 min read☕ Coffee break read

Original authors: Isabel Tsintsiper, Sheng Wong, Beth Albert, Shaun P Brennecke, Gabriel Davis Jones

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have four different "digital doctors" (AI chatbots) that are incredibly smart at reading medical textbooks and writing reports. You want to see if they are fair when diagnosing patients. To test this, the researchers created 50 fake patient stories. These stories were carefully written so that the patient's sex (male or female) didn't actually matter for figuring out what was wrong. It was like a puzzle where the solution was the same regardless of whether the piece was red or blue.

Here is what happened when they asked these AI doctors to solve the puzzles:

1. The "Guessing Game" Bias

When the researchers asked the AI, "What is the patient's sex?" without giving any clues, the AI didn't say, "I don't know." Instead, they started guessing based on stereotypes, like a person guessing someone's job based on their name.

  • The "Female" Guessers: Three of the AIs (ChatGPT, Claude, and DeepSeek) had a strong habit of guessing the patient was female. For example, if the story was about a heart issue or a urinary problem, they still often guessed "female."
  • The "Male" Guesser: One AI (Gemini) did the opposite, guessing male most of the time.
  • The Specialty Stereotypes: The bias got even stranger depending on the medical field.
    • If the story was about psychiatry, rheumatology, or blood disorders, all the AIs guessed "female" 100% of the time.
    • If the story was about heart or urinary issues, all the AIs guessed "male" 100% of the time.
    • It's as if the AIs have a mental map where "sadness" is always a woman and "heart trouble" is always a man, even when the story says nothing about that.

2. The "Temperature" Knob

The researchers tried turning a "temperature" dial on the AIs. Think of this like a spice level knob:

  • Low temperature (0.2): The AI is very strict, logical, and repetitive.
  • High temperature (1.0): The AI is more creative, random, and wild.

They hoped that making the AI more "creative" (turning up the heat) might break the stereotypes. It didn't. While the high-temperature AI gave slightly different lists of guesses, the underlying bias (guessing the wrong sex based on the specialty) stayed exactly the same. The "spice" didn't fix the recipe; it just made the same biased dish taste a little different.

3. The "I Don't Know" Option

In a second test, the researchers gave the AIs a third option: "Abstain" (or "I don't know").

  • ChatGPT immediately took the safe route and said "I don't know" every single time.
  • The others said "I don't know" most of the time, but not always.

The Catch: Even when the AIs stopped saying the patient's sex, they still thought about it. When asked to list possible diseases for a "male" version of the story versus a "female" version, the lists of diseases changed. The AI was still treating the two patients differently behind the scenes, even if it didn't explicitly label them.

4. The Bottom Line

The paper concludes that these AI models are like students who have memorized old, biased textbooks. They haven't learned to think neutrally; they are just repeating the patterns they saw in their training data.

  • The Bias is Built-In: The bias isn't a glitch; it's part of how the models are built.
  • Settings Don't Fix It: You can't just tweak a setting (like temperature) to make them fair.
  • The Danger: Because these models change their medical advice based on whether they think the patient is a man or a woman (even when the symptoms are identical), the authors say these general AI tools should not be used to make real medical decisions right now.

The Takeaway: If you ask these digital doctors for help, they might be smart, but they are also carrying heavy baggage of old-fashioned stereotypes. Until they are taught to ignore those stereotypes, humans need to keep a close eye on their work.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →