Medical Context Distorts Decisions in Clinical Vision Language Models
This paper reveals that clinical Vision-Language Models often fail in real-world scenarios by over-relying on textual information, being misled by irrelevant clinical history, and exhibiting high sensitivity to prompt variations, thereby underscoring the critical need for rigorous stress-testing and safeguards before their deployment in medical practice.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine a team of highly intelligent medical assistants (the AI models) whose job is to look at chest X-rays and decide if a patient is sick or healthy. They have two sources of information: the picture of the X-ray and the written notes from the doctor describing the patient's history.
The paper "Medical Context Distorts Decisions in Clinical Vision Language Models" acts like a stress test for these assistants. The researchers wanted to see if these AI "doctors" could actually trust the X-ray picture when the written notes said something different, or if they would just blindly follow the text.
Here is what they found, explained through simple analogies:
1. The "Text-Over-Image" Problem (The Loud Voice)
The Analogy: Imagine you are looking at a photo of a sunny beach. However, someone hands you a note that says, "This photo was taken during a blizzard." Even though your eyes clearly see the sun and sand, the AI assistants in this study mostly ignored the photo and believed the note.
The Finding: The researchers found that these AI models are text-heavy. When the X-ray image and the written report disagreed, the AI almost always sided with the text, even if the text was wrong.
- If the AI got the diagnosis right based on the picture, but the researchers swapped in a confusing or contradictory written report, the AI suddenly changed its mind and got it wrong.
- Surprisingly, if they swapped the picture for a different one but kept the original text, the AI barely noticed. It was like the picture was a whisper, and the text was a shout.
2. The "Irrelevant History" Trap (The Cluttered Desk)
The Analogy: Imagine a detective trying to solve a crime in a kitchen. But, someone keeps sliding notes onto their desk about crimes that happened in a different city, or about a different type of crime (like a broken leg instead of a theft). The detective gets confused and starts blaming the wrong person because of all the extra noise.
The Finding: In real hospitals, AI systems often pull up a patient's entire history, including old reports about knees, brains, or stomachs that have nothing to do with the current chest X-ray.
- The study showed that when the AI was fed these irrelevant past reports, its performance dropped. It got confused by the "noise."
- While the most advanced "frontier" models were better at ignoring this clutter, many of the open-source models got significantly worse as they were fed more and more unrelated history. They couldn't tell the difference between "important context" and "useless background noise."
3. The "Prompt Sensitivity" Issue (The Magic Wording)
The Analogy: Imagine asking a friend, "Is the sky blue?" They say "Yes." Then you ask, "Can you confirm the color of the sky?" They say "No." The question is the same, but the way you asked it changed their answer.
The Finding: The researchers asked the same medical question using four different styles of writing (e.g., a casual question, a formal doctor's order, a checklist).
- For the weaker AI models, changing the wording of the prompt caused them to flip their answers. One version of the question might say "Sick," and a slightly different version of the same question might say "Healthy."
- This means the AI's decision wasn't based on the medical facts, but on the specific phrasing of the request. This is dangerous because different doctors write notes in different styles; if the AI changes its mind just because the style changed, it's unreliable.
The Big Picture
The paper concludes that even though these AI models score very high on standard tests, they are fragile in real-world scenarios.
- They trust words more than pictures.
- They get confused by irrelevant history.
- They change their minds based on how you ask the question.
The authors argue that before we let these AI systems help make real medical decisions, we need to "stress-test" them much harder to ensure they don't get distracted by text, clutter, or phrasing. They need to learn to look at the X-ray first, rather than just reading the notes.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.