← Latest papers
💬 NLP

Same Patient, Different Words, Different Diagnosis? Evaluating Semantic Stability in Clinical LLMs

This paper proposes a semantic verification framework and new metrics to evaluate the stability of 16 clinical and general-purpose LLMs against meaning-preserving prompt variations, revealing that domain specialization does not consistently guarantee greater robustness in healthcare settings.

Original authors: Mahdi Alkaeed, Adnan Qayyum, Nabeel Abo Kashreef, Muhammad Bilal, Junaid Qadir

Published 2026-06-01
📖 5 min read🧠 Deep dive

Original authors: Mahdi Alkaeed, Adnan Qayyum, Nabeel Abo Kashreef, Muhammad Bilal, Junaid Qadir

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

The Big Idea: Same Patient, Different Words, Different Diagnosis?

Imagine you are a doctor. You have a patient with a specific set of symptoms. You write down a note describing the case. Then, you ask a computer program (an AI) to guess the diagnosis.

Now, imagine you rewrite that note. You use different words, change the sentence structure, or swap a medical term for a common one (like saying "heart attack" instead of "myocardial infarction"). The meaning is exactly the same. The patient hasn't changed.

The Problem: The paper asks: If you ask the AI the exact same question but phrase it differently, will it give you the same answer?

Unfortunately, the researchers found that these AI models are often like fickle weather forecasters. If you ask, "Will it rain?" you might get a "Yes." But if you ask, "Is there a chance of precipitation?" the same AI might suddenly say "No," even though the weather hasn't changed. In a hospital, this kind of inconsistency is dangerous.

The Solution: A "Truth Filter"

To test this properly, the researchers had to be very careful. They couldn't just ask the AI to rephrase things, because sometimes the AI accidentally changes the meaning (e.g., turning "no pain" into "pain").

They built a multi-layered truth filter to ensure the questions were truly identical in meaning:

  1. The Logic Check (NLI): They used a special "logic robot" that checks if two sentences logically imply each other. If Sentence A means Sentence B, and Sentence B means Sentence A, they are twins.
  2. The AI Judge: They used a second, very smart AI to double-check the logic robot's work.
  3. The Human Doctor: Finally, a real, licensed doctor reviewed the questions to make sure the medical facts (dosages, symptoms, timing) hadn't been accidentally altered.

Only the questions that passed all three tests were used for the experiment.

The Experiment: General vs. Specialized Doctors

The researchers tested 16 different AI models. They split them into two groups:

  • General-Purpose (GP): These are like smart generalists. They read everything on the internet and are good at many things, but they aren't medical experts by training.
  • Domain-Specific (DS): These are like specialized residents. They were trained specifically on medical books and patient records.

The Hypothesis: Everyone assumed the "Specialized Residents" (DS models) would be more stable and reliable than the "Generalists" (GP models) when the wording changed.

The Surprise: The specialized models did not consistently win.

  • Sometimes the specialized models were more stable.
  • Sometimes the general models were more stable.
  • Often, they were just as shaky as each other.

The Analogy: It's like testing two chefs. One is a master chef who only cooks Italian food (Specialized), and the other is a great cook who knows every cuisine (General). You ask them to make a pasta dish. If you describe the dish slightly differently ("boil the noodles" vs. "cook the pasta"), you might expect the Italian chef to be more consistent. But the study found that sometimes the Italian chef gets confused by the new words, while the general cook stays calm. Specialization alone doesn't guarantee stability.

The Confidence Trap: "I'm Sure, But I'm Wrong"

The researchers also looked at how confident the AI was in its answers.

  • The Finding: The AI often acts like a cocky student who is wrong but thinks they are right.
  • Even when the AI changed its answer because the wording changed (showing it was unstable), it often kept the same high level of confidence.
  • The Analogy: Imagine a student taking a test. If you rephrase a question, they might change their answer from "A" to "B." But if you ask, "How sure are you?" they still say, "100% sure!" The paper found that high confidence is not a good sign that the AI is actually correct or stable.

The Key Takeaways

  1. Small changes matter: Changing a few words in a medical prompt can cause the AI to give a completely different diagnosis, even if the meaning is the same.
  2. Training isn't a magic shield: Just because an AI is trained on medical data doesn't mean it will be more reliable when the language varies.
  3. Confidence is misleading: An AI saying it is "99% sure" doesn't mean it won't change its mind if you ask the question slightly differently.
  4. We need better testing: Before we trust these AIs in hospitals, we need to test them not just on whether they get the right answer once, but on whether they stay consistent when the words change.

In short: The paper warns us that current medical AIs are like unreliable translators. If you speak to them in "Medical English," they might give one answer. If you speak to them in "Casual English" (even if it means the same thing), they might give a different one. We need to fix this inconsistency before letting them help doctors treat patients.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →