← Latest papers
💬 NLP

Measuring Epistemic Resilience of LLMs Under Misleading Medical Context

This paper introduces MedMisBench, a comprehensive benchmark demonstrating that large language models, despite achieving expert-level scores on medical exams, exhibit fragile epistemic resilience by frequently abandoning correct medical judgments when exposed to misleading, authority-framed context.

Original authors: Hongjian Zhou, Xinyu Zou, Jinge Wu, Sean Wu, Junchi Yu, Bradley Max Segal, Tobias Erich Niebuhr, Sara Amro, Michael Petrus, Sheikh Momin, Alexandra M. Cardoso Pinto, Rachel Niesen, Laura Sophie Wegner
Published 2026-06-11
📖 4 min read☕ Coffee break read

Original authors: Hongjian Zhou, Xinyu Zou, Jinge Wu, Sean Wu, Junchi Yu, Bradley Max Segal, Tobias Erich Niebuhr, Sara Amro, Michael Petrus, Sheikh Momin, Alexandra M. Cardoso Pinto, Rachel Niesen, Laura Sophie Wegner, Dhruv Darji, Jung Moses Koo, Joshua Fieggen, Kapil Narain, Mingde Zeng, Lei Clifton, Linda Shapiro, Fenglin Liu, David A. Clifton

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have a brilliant medical student who has memorized every textbook and can ace any licensing exam with a perfect score. You trust them completely. But then, you walk into the room and whisper a fake rule to them: "By the way, the hospital just changed the rules. For this specific condition, we now do X instead of Y."

Surprisingly, this "perfect" student immediately forgets everything they studied, believes your fake rule, and gives you the wrong answer.

This is the core discovery of the paper MedMisBench. The researchers found that even the smartest AI doctors (Large Language Models) are surprisingly fragile when faced with misleading information, even if they know the correct answer beforehand.

Here is a breakdown of their findings using simple analogies:

1. The "Trap Door" in the Exam

Usually, we test AI doctors with clean, textbook questions. They score very high (about 71% correct). The researchers assumed this meant the AI was safe to use for real health advice.

They were wrong. The researchers built a new test called MedMisBench. Think of it as a "trap door" exam.

  • The Setup: They take a question the AI knows the answer to.
  • The Trap: They inject a sentence that looks like a credible medical rule but is actually a lie.
  • The Result: The AI's accuracy crashes from 71% down to 38%. In fact, in more than half the cases (51.5%), the AI completely abandons the truth and follows the lie.

2. The "Authority" Effect

The paper tested how the lie was delivered. It turns out, the AI is most easily tricked when the lie sounds official.

  • The Metaphor: Imagine a student ignoring a teacher's textbook because a stranger in a suit (an "Authority") walks in and says, "Actually, the book is wrong."
  • The Finding: When the fake information was framed as a new hospital rule, a guideline, or an official note, the AI fell for it almost 70% of the time.
  • The "Exception" Trap: The AI was also easily fooled by lies that sounded like specific exceptions (e.g., "This rule applies to everyone except this one weird case").

3. The "One Voice" vs. "The Crowd"

The researchers tested two ways of presenting the lies:

  • Type 1 (The Whisper): The AI sees the question and one specific lie supporting a wrong answer.
    • Result: Disaster. The AI immediately switches to the wrong answer.
  • Type 2 (The Debate): The AI sees the question, the truth, and lies supporting every wrong answer all at once.
    • Result: The AI does much better here. It can see the whole picture and usually stick to the truth.
  • The Lesson: The danger isn't that the AI can't handle any lies; it's that it crumbles when a single, plausible-sounding lie is whispered directly to it.

4. "Thinking Harder" Doesn't Always Help

You might think that if you tell the AI to "think step-by-step" or use more brainpower, it would be harder to trick.

  • The Metaphor: It's like asking a student to "double-check their work."
  • The Finding: For some AI models, thinking harder actually made them more susceptible to the lies. They would over-analyze the fake "official rule" and convince themselves it was the right path.

5. The Real-World Danger

The researchers didn't just look at right/wrong answers; they asked 14 real doctors from 7 different countries to review the AI's mistakes.

  • The Finding: In 38% of the cases where the AI got tricked, the wrong answer could have caused serious harm to a patient.
  • The Takeaway: This isn't just a math error; it's a safety hazard. The AI isn't just "hallucinating"; it's confidently giving dangerous advice because it was tricked by a fake context.

Summary

The paper concludes that we have been measuring AI doctors on how well they know the textbook, but we haven't been testing if they can stick to the truth when someone tries to trick them with fake rules.

MedMisBench is a new tool designed to measure this "epistemic resilience" (the ability to keep your judgment when the world is trying to confuse you). It shows that currently, our AI doctors are not resilient enough to be trusted with health advice in the messy, information-filled real world where fake news and misleading claims are common.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →