← Latest papers
💻 computer science

Contextual Cue Susceptibility in Clinical AI: A Randomized Controlled Clinician-Comparator Study

This randomized controlled study reveals that large language models are significantly more susceptible than human clinicians to contextual cues that induce incorrect diagnostic choices, highlighting a critical safety vulnerability for clinical AI systems that can be mitigated through specific prompting strategies.

Original authors: Mahmud Omar, Reem Agbareia, Alexander Charney, Raja-Elie Abdulnour, Ankit Sakhuja, Eyal Klang, Girish Nadkarni, Benjamin Glicksberg

Published 2026-07-29
📖 4 min read☕ Coffee break read

Original authors: Mahmud Omar, Reem Agbareia, Alexander Charney, Raja-Elie Abdulnour, Ankit Sakhuja, Eyal Klang, Girish Nadkarni, Benjamin Glicksberg

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are a detective trying to solve a mystery. You have a list of clues, and your job is to figure out who did it. Sometimes, a clue is huge and obvious, like a muddy boot print. Other times, a clue is tiny and weird, like a single blue thread found on the suspect's coat. In the world of medicine, doctors are the detectives, and their "clues" are the symptoms patients describe. For a long time, we've known that human detectives can get tricked. If a story mentions a rare, scary detail, a doctor might get distracted and forget that the most common illness is still the most likely answer. This is called "availability bias"—our brains latch onto the vivid story instead of the boring statistics.

Now, imagine we give our detective a super-smart robot assistant. This robot is a Large Language Model (LLM), a type of artificial intelligence that reads millions of books and learns to talk and think like a human. These robots are getting great at summarizing medical records and suggesting diagnoses. But here's the big question: Is the robot just as good at ignoring the tiny, distracting clues as a human is? Or is the robot even more easily tricked? Scientists are worried because if a robot gets distracted by a tiny detail, it could make thousands of mistakes very quickly, especially if it starts acting on its own to book appointments or write notes without a human checking its work.

This paper sets up a giant, controlled experiment to find out. The researchers created 100 short medical stories, like mini-mysteries. In some versions of the story, they added a tiny, harmless detail—like a patient mentioning they recently traveled to a specific country or worked in a specific job. This detail pointed toward a rare, "lure" diagnosis that sounded exciting but was actually unlikely. In the other versions, that tiny detail was missing. Then, they asked two groups to solve the cases: 47 human medical professionals (doctors, residents, fellows, and medical students) and 22 different AI models. They wanted to see if the tiny detail would change the answer, and if so, who would be tricked more: the humans or the robots.

The results were a bit shocking. When the tiny "travel" or "job" clue was added, the human doctors were still influenced by it, choosing the rare, wrong "lure" diagnosis about 28% of the time, though they stuck with the correct diagnosis about 55% of the time. The AI models, however, fell for the trick like a magic show. When the clue was there, the AI models chose the rare, wrong "lure" diagnosis a massive 85% of the time! That means the AI was tricked by that single tiny sentence about 57 percentage points more often than the humans were. The AI didn't just get confused; it confidently convinced itself that the rare disease was the answer, even though the math said it shouldn't be.

The researchers also tested if they could "fix" the robot by changing the instructions they gave it. When they told the AI, "Remember that common things are common," the AI stopped getting tricked so much, and its performance looked more like the humans. But when they told the AI, "Be on the lookout for rare and serious things," the AI went crazy, getting tricked even more often, choosing the wrong answer 93% of the time. This suggests the AI isn't "thinking" in a stable way; it's just reacting wildly to the words in the prompt.

The main takeaway isn't that the AI is bad at medicine in general—it can actually be very accurate when the clues are clear. The problem is that these models are incredibly sensitive to context. A tiny, clinically unimportant sentence can completely flip their decision, much more than it flips a human doctor's mind. The authors warn that as these AI systems start doing more work on their own—like writing notes or triaging patients without a human looking over their shoulder—this sensitivity to tiny cues could become a serious safety risk. If a robot changes its mind just because a patient mentioned a vacation, that's a problem we need to solve before we let the robots drive the car.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →