Gender-Dependent Diagnostic Substitution in LLM Medical Triage: Same Symptoms, Unequal Urgency
This study reveals that large language models exhibit a systemic gender bias in medical triage by disproportionately downgrading emergency referrals for young women with severe neurological symptoms compared to men, a disparity driven by the models' reliance on gendered diagnostic priors that substitute urgent conditions with less critical, epidemiologically associated diagnoses.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you have three very smart, highly trained digital doctors (AI models). You ask them the exact same question: "I have a constant headache, my vision is blurry, I'm nauseous in the morning, and I see spots."
Now, here is the twist: You tell the AI that the patient is a 25-year-old man. Then, you ask the exact same question again, but this time you say the patient is a 25-year-old woman.
According to this paper, the AI gives you two completely different answers, even though the symptoms are identical.
The "Traffic Light" Problem
Think of medical triage (deciding how urgent a problem is) like a traffic light system.
- Red Light: Go to the Emergency Room (ER) immediately.
- Yellow Light: See a regular doctor soon.
- Green Light: Take some rest and wait.
The paper found that when the AI thought the patient was a young man, it almost always flashed a Red Light. It said, "This is dangerous, go to the ER right now!"
But when the AI thought the patient was a young woman with the exact same symptoms, it often flashed a Yellow Light. It said, "This is likely a specific condition called IIH (Idiopathic Intracranial Hypertension), which is common in young women. You should make an appointment with a regular doctor, but you don't need the ER."
The "Shortcut" Trap (Diagnostic Substitution)
Why did the AI do this? The paper calls it "Diagnostic Substitution."
Imagine the AI is a detective who has read millions of medical textbooks. It knows a statistical fact: "Headaches and vision problems in young women are often caused by a condition called IIH."
When the AI sees a young woman, it takes a shortcut. It immediately locks onto that specific diagnosis (IIH). Because IIH is usually treated by a specialist in an office rather than an emergency room, the AI decides the situation isn't an emergency.
However, when the AI sees a young man, it doesn't lock onto IIH as quickly. Instead, it keeps its options wide open, thinking, "This could be a tumor or a blockage." Because those possibilities are scary and need immediate checking, it sends the man to the ER.
The Flaw: The paper argues that the AI made a mistake in logic. Even if the woman does have IIH, the symptoms described (severe headache, vision loss, nausea) are red flags that require urgent attention regardless of the cause. By guessing the diagnosis too early, the AI accidentally turned a "Red Light" into a "Yellow Light" for women.
The "Age" Switch
Here is the most interesting part: The bias disappears as the patient gets older.
When the paper tested the AI with 65-year-olds (both men and women), the AI stopped making this distinction. It sent almost everyone to the ER.
Why? Because the statistical shortcut (IIH) mostly affects young women of childbearing age. Once the AI sees a 65-year-old, that shortcut no longer applies, so it treats men and women the same way again. This proves the AI wasn't just "hating" women; it was blindly following a statistical rule that didn't fit the specific situation.
The Bottom Line
The paper concludes that these AI models aren't being "sexist" in a mean-spirited way. Instead, they are being too clever for their own good. They learned a real medical pattern (that young women often get IIH) and applied it too rigidly.
In doing so, they prioritized the most likely diagnosis over the most urgent action. The paper warns that for medical AI, we need to teach the system: "Even if you think you know what the disease is, if the symptoms look dangerous, treat it as an emergency first."
In short: The AI saw a young woman, thought, "Oh, that's just a common thing for women," and sent her home. It saw a young man, thought, "That could be anything dangerous," and sent him to the ER. The paper shows that for these specific symptoms, both should have been sent to the ER.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.