← Latest papers
💬 NLP

Same Facts, Different Diagnosis: Measuring and Mitigating Narrative Anchoring in Clinical Language Models

This paper identifies "Narrative Anchoring," a bias where clinical language models produce divergent diagnoses for identical facts based solely on sociolinguistic register, and proposes "NarrativeShield," a three-agent pipeline that structurally extracts and verifies facts to nearly eliminate this gap while maintaining diagnostic stability.

Original authors: Prabhjot Singh, Pritam Deka, Vijay Chennareddy

Published 2026-07-31
📖 4 min read☕ Coffee break read

Original authors: Prabhjot Singh, Pritam Deka, Vijay Chennareddy

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are a detective trying to solve a mystery, but instead of looking at the clues, you are judging the suspect based on how they tell their story. If they speak with a fancy, polished voice, you might trust them more. If they speak with a rough, working-class accent or use colorful metaphors, you might doubt them, even if the actual facts of the crime are exactly the same. This is the world of "Large Language Models" (LLMs), which are super-smart computer programs trained to read and write like humans. Scientists are now trying to use these programs as "digital doctors" to help diagnose illnesses. But there's a scary problem: these digital doctors might be biased not just by who the patient is (like their race or income), but by how the patient describes their pain. This paper dives into that hidden bias, asking a simple but profound question: If two patients have the exact same symptoms, will the computer give them the same diagnosis if one tells the story in a "rich" way and the other in a "struggling" way?

The researchers call this sneaky problem Narrative Anchoring. It's like a compass that gets stuck pointing toward the style of the story rather than the truth of the facts. To test this, they didn't just ask the computer to guess; they built a massive, super-strict game. They took 1,000 real medical exam questions (the kind doctors have to pass to get their licenses) and rewrote each one three times. One version was neutral, one was rewritten to sound like someone facing financial stress, and a third sounded like someone using cultural metaphors. Crucially, they made sure the medical facts—like the fever, the pain location, and the blood test results—stayed exactly the same in all three versions. They even hired a second computer and real human doctors to double-check that no facts were lost or changed during the rewriting.

When they fed these stories to seven different AI models, the results were startling. Even though the medical facts were identical, the AI models gave different diagnoses depending on the "voice" of the story. This happened in every single model they tested. The gap in accuracy was significant: for some models, the "neutral" voice got the right answer about 15% more often than the "struggling" voice. The researchers tried to fix this by asking the AI to "think step-by-step" (a method called Chain-of-Thought) or by simply telling it, "Ignore the style, focus on the facts." These tricks helped a little, but not enough. In fact, asking the AI to think harder sometimes made it worse at getting the right answer overall, like a student who over-analyzes a simple math problem and gets it wrong.

So, the team built a new tool called NarrativeShield. Think of it as a strict editor who sits between the patient and the doctor. Before the AI doctor ever sees the story, this editor strips away all the fancy words, the emotional tone, and the cultural metaphors, leaving only a bare-bones list of cold, hard medical facts. Then, the AI doctor only looks at that list. When they used NarrativeShield, the bias almost vanished. The gap between the different voices dropped to nearly zero (between -0.004 and 0.037), meaning the AI started treating all stories fairly, regardless of how they were told. However, there was a catch: for most models, this strict editing made the AI slightly less accurate overall, like a translator who removes all the flavor from a poem to ensure the meaning is perfect.

The paper also tested a "bare-bones" medical AI that hadn't been taught how to follow instructions well. This model failed completely when asked to follow the new rules, proving that you can't just tell a computer to "be fair" if it doesn't have the basic ability to listen to instructions in the first place. The authors conclude that the problem isn't just about what labels we give patients; it's about how the computer processes the way patients speak. To fix it, we can't just change the prompt; we have to change the architecture of the system itself, ensuring the "facts" are separated from the "story" before the diagnosis is made. While this is a promising step, the researchers are careful to say this is still a research tool, not a ready-made product for real hospitals yet, and more testing is needed to see if it works in the messy, real world of patient-doctor conversations.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →