← Latest papers
💬 NLP

Counterfactual Cultural Cues Reduce Medical QA Accuracy in LLMs: Identifier vs Context Effects

This paper introduces a counterfactual benchmark demonstrating that inserting non-decisive cultural identifiers and context into medical questions significantly degrades the diagnostic accuracy of various large language models, often causing clinically correct reasoning to fail when cultural cues are present.

Original authors: Amirhossein Haji Mohammad Rezaei, Zahra Shakeri

Published 2026-01-29
📖 5 min read🧠 Deep dive

Original authors: Amirhossein Haji Mohammad Rezaei, Zahra Shakeri

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have a brilliant, super-smart AI doctor. This doctor has read every medical textbook ever written and can diagnose diseases with incredible speed. But there's a catch: this doctor is a bit too sensitive to the "flavor" of a patient's story, even when that flavor doesn't actually change the medical facts.

This paper is like a stress test for that AI doctor. The researchers wanted to see if the AI would give a different diagnosis just because they changed the patient's cultural background, even though the symptoms and the correct answer stayed exactly the same.

Here is the story of what they found, broken down into simple parts:

1. The Setup: The "Same Cake, Different Wrapper"

The researchers took 150 real medical exam questions (like the ones doctors take to get licensed). These questions describe a patient with specific symptoms, and there is only one correct medical answer.

Then, they created 1,650 new versions of these questions. They didn't change the symptoms or the disease. Instead, they just added "cultural sprinkles" to the story. They did this in three ways:

  • The ID Tag: They added a sentence saying who the patient is (e.g., "This is a 35-year-old Muslim man from the Middle East").
  • The Context: They added a sentence about what the patient was doing (e.g., "The pain started while he was praying at the mosque").
  • The Combo: They added both the ID and the Context.

They did this for three specific groups: Indigenous Canadians, Middle-Eastern Muslims, and Southeast Asians. They also made "neutral" versions (just adding a boring sentence like "The patient arrived with a family member") to make sure the AI wasn't just getting confused by the extra words.

The Golden Rule: A real human doctor reviewed every single new version and confirmed: "The correct answer hasn't changed. The disease is still the same."

2. The Experiment: Testing the AI Doctors

They fed these questions to five different AI models (including big names like GPT-5.2, Llama, and MedGemma). They asked the AI to pick the right answer, sometimes just giving the letter (A, B, C) and sometimes asking for a short explanation first.

Think of it like asking a student to solve a math problem. If you tell the student, "This problem is for a student named Ahmed," does the student suddenly get the math wrong?

3. The Results: The AI Got Confused by the "Flavor"

The results were surprising and a bit worrying.

  • The "Combo" Effect: When the AI saw both the patient's identity and the cultural context together, it got the most confused. The accuracy dropped significantly (by about 3 to 7 percentage points). It's like the AI started thinking, "Oh, this is a Muslim man praying? Maybe the answer is different for him," even though the medical facts said otherwise.
  • The "ID" Problem: Surprisingly, just adding the patient's identity (the "ID Tag") caused almost as much trouble as the full combo. The AI seemed to latch onto the label (e.g., "Muslim," "Indigenous") and let that label change its reasoning.
  • The "Context" Problem: Changing just the context (what they were doing) had a smaller effect, but it still messed things up.
  • The "Neutral" Test: When they added boring, neutral sentences, the AI's performance stayed mostly the same. This proved that the AI wasn't just getting tired of reading long sentences; it was specifically reacting to the cultural words.

4. The "Why": The AI's Bad Habits

The researchers looked at why the AI was failing. They found that when the AI got the answer wrong, its explanation often sounded like it was making up stereotypes.

  • The "Guessing Game": Instead of looking at the symptoms, the AI started guessing based on cultural assumptions. For example, it might assume a certain group eats a specific diet or has a specific lifestyle, and then use that guess to pick the wrong disease.
  • The "Flip": If the AI got the answer right on the original question, adding cultural cues often made it "flip" its answer to the wrong one. It's like a student who knows the answer is "4," but then sees a picture of a cat in the corner of the page and suddenly thinks, "Oh, maybe it's 5 because cats have 5 lives?"

5. The Big Takeaway

The paper concludes that current medical AI models are not culturally neutral. They are like a chef who changes the recipe just because the customer is wearing a specific hat.

Even though the medical facts didn't change, the AI's "diagnosis" changed based on cultural clues. This means that if we use these AI tools in real hospitals without fixing this problem, they might give different (and potentially dangerous) advice to patients just because of their background.

In short: The AI is smart, but it's also a bit prejudiced. It lets cultural labels cloud its medical judgment, turning a simple diagnosis into a guessing game based on stereotypes rather than science. The researchers released their test questions so others can try to fix this "cultural blindness" in future AI models.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →