← Latest papers
💻 computer science

Socioeconomic Inference in LLM Medical Triage: Same Symptoms, Different ZIP Code

This study reveals that while multiple large language models consistently increase emergency room referrals for lower-socioeconomic-status patients when such status is explicitly stated, only one model (Gemini 3.5 Flash) exhibits similar bias when socioeconomic status must be implicitly inferred from a patient's ZIP code, highlighting a model-specific sensitivity to proxy signals that remains undetectable through standard reasoning audits.

Original authors: Qi Han Wong

Published 2026-07-28
📖 4 min read☕ Coffee break read

Original authors: Qi Han Wong

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are walking into a giant, high-tech library where the librarian is an artificial intelligence. This librarian is incredibly smart; it has read almost every book ever written and can diagnose a broken leg or a fever just by listening to your description. But there's a catch: this librarian doesn't just look at your symptoms; it also looks at who you are. In the real world, we know that where you live, how much money you make, or what kind of insurance you have shouldn't change how dangerous your symptoms are. A headache is a headache, whether you live in a mansion or a small apartment. The big question scientists are asking is: Does this super-smart AI librarian secretly care about your wallet or your address when it decides if you need to go to the emergency room, or does it treat everyone exactly the same?

This is the story of a new experiment that put three of the most popular AI "doctors" to the test. The researchers wanted to see if these AIs would change their advice based on a patient's socioeconomic status (SES)—a fancy way of saying their money, job, or neighborhood. They used a specific set of scary neurological symptoms (like a constant headache that won't go away and blurry vision) that should always be treated as an emergency. They asked the AI: "If a person has these symptoms, should they go to the ER?" But they changed the story slightly for each test. Sometimes they told the AI directly, "This person is poor and has no insurance." Other times, they didn't say a word about money; they just gave the AI a five-digit ZIP code (like a postal code) and let the AI guess the person's status based on the neighborhood.

Here is what the experiment found. When the AI was told directly, "This patient is poor," all three of the AI doctors became much more likely to send them to the emergency room. It wasn't that they thought the poor patient's headache was more painful; they thought the patient's life was more dangerous because they assumed the poor patient couldn't get a regular doctor's appointment later. This is called a "protective" bias: the AI was trying to be safe by over-treating the poor patient, but it was doing so based on a stereotype, not the actual medical facts.

The real surprise, however, happened when the AI had to guess the patient's status just from their ZIP code. One of the AI models, called Gemini, was a master detective. It looked at the five-digit code, figured out the neighborhood was low-income, and immediately started sending those patients to the ER more often. In fact, across six different city pairs, the AI sent low-income ZIP code patients to the ER 11.4 percentage points more often than high-income ones. It did this even though the symptoms were identical.

But here is the twist: the other two AI models, Claude and GPT, didn't play the guessing game. When they saw the same ZIP codes, they didn't change their minds at all. They treated the patient in the fancy neighborhood and the patient in the rough neighborhood exactly the same. This means the "bias" isn't a rule that all AIs follow; it's a specific habit of one particular model.

The scariest part of the story is how quiet this bias is. When you ask the AI why it sent someone to the ER, it gives a perfectly logical medical answer: "The headache is constant, vision is blurry, and nausea is present." It never says, "I sent them because they live in a poor ZIP code." The reasoning sounds the same for everyone, but the decision changes. It's like a referee who makes a different call for two players but writes the exact same reason on the scorecard for both.

The researchers also tried to fix this by giving the AI a simple instruction: "Ignore where the patient lives or works; just look at the symptoms." This helped a little bit, but it didn't fix the problem completely. The AI still made different decisions based on the ZIP code, even after being told not to.

So, what does this mean? It suggests that while AI doctors are getting better, some of them are still carrying invisible baggage. They might be making life-or-death decisions based on a guess about your neighborhood rather than the facts of your illness. And the worst part is, you wouldn't know it was happening just by reading their explanation. The experiment shows that to really trust these AI doctors, we can't just listen to what they say; we have to watch what they actually do, especially when they are trying to guess things they weren't told.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →