← Latest papers
💬 NLP

When Large Language Models Fail in Healthcare: Evaluating Sensitivity to Prompt Variations

This study systematically evaluates the sensitivity of general-purpose and medical-specific Large Language Models to prompt variations using the MedMCQA benchmark, revealing that even minor lexical or syntactic perturbations can significantly degrade their reliability and induce harmful clinical outputs, thereby highlighting critical safety risks for their deployment in healthcare.

Original authors: Mahdi Alkaeed

Published 2026-06-08
📖 4 min read☕ Coffee break read

Original authors: Mahdi Alkaeed

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

The Big Idea: The "Fickle Chef" of Medicine

Imagine you have a world-class chef (the Large Language Model, or LLM) who is supposed to cook the perfect meal for a patient based on a recipe card (the prompt). This chef is incredibly smart and knows thousands of recipes.

However, this paper discovers a scary flaw: This chef is extremely sensitive to how the recipe is written.

If you change a single word, rearrange the order of the ingredients, or even just spell a word slightly differently, the chef might suddenly decide to serve a completely different dish—or worse, a poisonous one. The paper argues that in healthcare, where a "dish" is a medical diagnosis or advice, this kind of unpredictability is dangerous.

The Experiment: Testing the Chef's Stability

The researchers wanted to see if these medical AI chefs could handle small changes without messing up. They used a massive library of medical questions (called MedMCQA) and tested two types of "chefs":

  1. General Chefs: Smart AI models not specifically trained for medicine (like GPT-3.5 or Llama3).
  2. Specialist Chefs: AI models trained specifically on medical books and journals (like BioBERT or ClinicalBERT).

They tested the chefs under three conditions:

  • The Original Recipe: The standard, clean question.
  • The "Synonym Swap" (Lexical Perturbation): Changing words to their synonyms (e.g., changing "shortness of breath" to "dyspnea") or fixing typos.
  • The "Reordered Kitchen" (Syntactic Perturbation): Changing the structure of the sentence (e.g., switching from "The nurse gave the drug" to "The drug was given by the nurse").

What They Found: The "Fragile" Truth

The results showed that no chef is truly safe yet. Here is what happened:

1. The "Synonym Swap" was mostly okay (but not perfect)
When the researchers just swapped words for similar ones (like changing "five mg" to "5 mg"), most chefs stayed calm. They still gave the right answer. It's like if you asked for "a glass of water" instead of "a tumbler of water," the chef still brings water.

  • However, even here, the chefs sometimes got confused about how to present the answer (like forgetting to write it in a specific format), even if the medical advice was correct.

2. The "Reordered Kitchen" caused chaos
When the researchers changed the sentence structure or the order of the options, the chefs started to fail.

  • The Analogy: Imagine a chef who is great at following a list, but if you put the last item on the list first, they forget the first item entirely.
  • The Result: In some cases, simply rephrasing the question caused the AI to give a wrong diagnosis. For example, a patient with symptoms of COPD might be diagnosed with Pulmonary Edema just because the sentence structure was tweaked.

3. The "Specialist Chefs" weren't immune
You might think a chef trained only on medical books would be safer. The paper found that specialist chefs (BioBERT, etc.) were actually just as fragile, and sometimes more sensitive to structural changes, than the general chefs. They didn't have a "superpower" against these tricks.

The "Adversarial" Danger: The Sneaky Note

The paper also looked at "adversarial" prompts—these are like someone slipping a sneaky note into the recipe card to trick the chef.

  • The Analogy: Imagine someone whispering, "Ignore the first step, the patient is actually allergic to this," even though the patient isn't.
  • The Result: These small, tricky changes could make the AI recommend the wrong dosage of medicine or miss a critical symptom entirely. The paper calls this "clinically dangerous."

The Verdict: Why This Matters

The paper concludes that current medical AI is not "intrinsically safe."

  • The Problem: If a doctor asks the AI the same question in two different ways, the AI might give two different answers. One might be right, and the other might be wrong.
  • The Risk: In a hospital, you can't have a system that changes its mind just because you rephrased a sentence. If the AI hallucinates a drug interaction or misses a diagnosis because of a typo, patients could get hurt.

Summary in a Nutshell

Think of these AI models as brilliant but nervous students. They know the material perfectly when the test is written clearly. But if the teacher changes the font, swaps a few words, or rearranges the paragraphs, the student might panic and get the answer wrong.

The paper says: "We cannot trust these students to grade medical exams yet because they are too easily confused by how the question is asked." Before we let them help doctors, we need to teach them to be more stable and less sensitive to the way we talk to them.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →