Calibrated? Not for Everyone: How Sexual Orientation and Religious Markers Distort LLM Accuracy and Confidence in Medical QA
This paper reveals that social identity markers, particularly sexual orientation and religious affiliation, significantly distort both the accuracy and uncertainty calibration of medical Large Language Models, creating a "calibration crisis" that poses serious risks to equitable and safe clinical deployment.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you have a super-smart medical student named "LLM" (Large Language Model). This student has read every medical textbook in the world and can diagnose diseases with incredible speed. Doctors are starting to trust this student to help them make decisions.
But there's a catch: For a medical student to be truly safe, they need to know when they are unsure. If they are 99% sure they are right, they should be right. If they are only 50% sure, they should say, "I don't know, ask a human doctor." This is called being "calibrated."
This paper is like a detective story that asks: What happens when we tell this medical student about the patient's personal life, like their religion or who they love?
The Experiment: The "Identity" Test
The researchers took thousands of medical questions (like "A 45-year-old woman has chest pain; what is wrong?") and created three versions of each:
- The Neutral Version: Just the medical facts.
- The "Heterosexual" Version: Added a sentence saying, "The patient identifies as heterosexual."
- The "Homosexual" Version: Added a sentence saying, "The patient identifies as homosexual."
- The "Intersectional" Version: Added both sexual orientation and religion (e.g., "The patient is homosexual and Catholic").
They ran these questions through 9 different AI models to see if the answers changed.
The Big Discovery: The "Confidence Trap"
Here is the scary part, explained with a simple analogy:
Imagine the AI is a weather forecaster.
- Scenario A (Neutral): You ask, "Will it rain?" The AI says, "Yes, 90% chance." It rains. Perfect.
- Scenario B (The Bias): You ask the same question, but you mention, "I am a gay man." Suddenly, the AI says, "Yes, 90% chance," but it doesn't rain.
The AI got the answer wrong, but it was just as confident as when it was right.
In the paper, they found that when they added "homosexual" or specific religious labels:
- Accuracy Dropped: The AI started giving the wrong medical answers more often.
- Confidence Stayed High (or got worse): The AI didn't realize it was confused. It kept shouting, "I'm sure!" even when it was wrong.
- The "Double Trouble" Effect: When they combined identities (e.g., "Homosexual and Muslim"), the AI got even more confused and more wrong than if you just added one label. It's like the AI's brain started mixing up stereotypes instead of looking at the medical facts.
Why This is Dangerous
Think of the AI as a co-pilot for a doctor.
- If the co-pilot says, "I'm 90% sure this is a heart attack," the doctor might rush to treat it.
- If the co-pilot is biased and says, "I'm 90% sure this is a heart attack" (when it's actually something else) just because the patient is gay, the doctor might treat the wrong thing.
- The Worst Part: Usually, if an AI is unsure, it says, "Hey, I'm not sure, let's ask a human." But because the AI's "confidence meter" is broken by these identity labels, it fails to ask for help. It stays silent and confident while making a mistake.
The "Open-Ended" Test
The researchers worried that maybe the AI was just bad at picking multiple-choice answers. So, they asked the AI to write the answer in its own words (like a real doctor writing a report).
- Result: The problem got worse.
- Example: In one case, a patient had liver failure. The AI correctly identified it when the patient was neutral. But when the text said the patient was "homosexual," the AI suddenly guessed "HIV-related kidney disease" (a stereotype) instead of the liver issue, even though the medical symptoms were identical. And it was very confident about this wrong guess.
The Takeaway
The paper concludes that social identity markers act like "noise" that breaks the AI's internal compass.
- It's not just about being "politically incorrect." It's about safety.
- If an AI cannot be trusted to give the same level of confidence for a gay patient as it does for a straight patient, we cannot safely use it in hospitals.
- Simply deleting these words from the patient's file isn't a fix, because the AI can often guess these things from other clues (like how the story is written).
In short: We are building AI doctors that are smart but easily distracted by who the patient is. Until we fix their "confidence meter" so it works for everyone, regardless of their identity, we can't let them drive the car alone.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.