Consistent but Dangerous: Per-Sample Safety Classification Reveals False Reliability in Medical Vision-Language Models
This paper demonstrates that consistency under paraphrase is a flawed proxy for reliability in medical vision-language models because models can achieve perfect consistency by ignoring images and relying on text patterns, a dangerous failure mode that standard confidence metrics miss but can be exposed by comparing predictions against a text-only baseline.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you hire a new assistant to help you diagnose medical problems based on X-ray images. You want this assistant to be reliable. You test them by asking the same question in different ways:
- "Is there fluid in the lungs?"
- "Can you see fluid in the lungs?"
- "Does the patient have a pleural effusion?"
If your assistant gives the exact same answer every time, no matter how you phrase the question, you think, "Great! They are consistent. They are reliable."
This paper argues that this feeling of safety is a dangerous illusion.
Here is the simple breakdown of what the researchers found, using some everyday analogies.
1. The "Stubborn Chef" Analogy
Imagine a chef who claims to cook based on the ingredients you give them.
- The Test: You ask, "Is this soup salty?"
- The Problem: The chef always answers "Yes," no matter what soup you show them. Even if you show them a bowl of plain water, they say "Yes." Even if you ask, "Is this soup salty?" in French, German, or a different accent, they still say "Yes."
The Trap:
- Consistency: The chef is 100% consistent. They never change their answer based on how you ask.
- Reliability: You might think, "Wow, they are so confident and stable!"
- The Reality: The chef isn't looking at the soup at all. They are just guessing "Yes" because they know that, statistically, most soups in this restaurant are salty. They are ignoring the actual image (the soup) and relying on a text shortcut.
In the world of AI, this is called the "Dangerous" quadrant. The AI is consistent, but it's not actually looking at the X-ray.
2. The Four Types of AI Assistants
The researchers created a simple chart to sort AI assistants into four groups based on two questions:
- Are they consistent? (Do they give the same answer if you rephrase the question?)
- Are they looking at the picture? (If you take the picture away, do they change their answer?)
Here are the four types:
🌟 The Ideal Assistant (The Doctor):
- Behavior: They give the same answer no matter how you ask, AND they change their answer if you swap the X-ray for a blank piece of paper.
- Verdict: Safe. They are looking at the image and thinking clearly.
🥚 The Fragile Assistant (The Nervous Intern):
- Behavior: They look at the picture, but if you change the wording of the question slightly, they get confused and give a different answer.
- Verdict: Risky, but honest. They are trying to use the image, but they are easily confused by language.
☠️ The Dangerous Assistant (The Stubborn Chef):
- Behavior: They give the exact same answer no matter how you ask, BUT if you take the picture away, they give the exact same answer.
- Verdict: DEADLY. They look super reliable because they are consistent, but they are actually ignoring the patient's X-ray entirely. They are just guessing based on patterns in the text.
- The Twist: These assistants often get high scores on accuracy tests! If the disease is common, guessing "Yes" all the time gets you a high score, even if you never looked at the X-ray.
🗑️ The Worst Assistant (The Broken Robot):
- Behavior: They change their answer when you rephrase the question, AND they ignore the picture.
- Verdict: Useless. They are inconsistent and not looking at the image.
3. The Big Surprise: "Fixing" Made It Worse
The researchers tried to "fix" the AI assistants by training them to be more consistent (so they wouldn't get confused by different wordings).
The Result?
The training worked! The assistants became very consistent. But, it pushed them from the "Fragile" group into the "Dangerous" group.
- They stopped getting confused by wording.
- Instead, they started ignoring the X-rays entirely and just guessing the most common answer.
It's like teaching a student to pass a test by memorizing the answer key instead of learning the subject. They get a perfect score (high consistency), but they can't actually do the work (they aren't looking at the image).
4. Why Standard Tests Failed
Usually, when we test AI, we check:
- Accuracy: Did they get the right answer? (The "Dangerous" AI often gets this right by luck).
- Confidence: How sure are they? (The "Dangerous" AI is very confident because they have a simple rule to follow).
- Consistency: Do they answer the same way if you rephrase? (The "Dangerous" AI is perfect here).
The paper says: If you only check these three things, you will hire the "Dangerous" assistant and fire the "Ideal" one. The "Dangerous" AI looks perfect on paper but is actually blind.
5. The Simple Solution: The "No-Picture" Test
The researchers propose a very simple, cheap test to catch these dangerous AI models before we use them in hospitals.
The Test:
Run the AI's question through the system twice:
- Once with the X-ray image.
- Once without the image (just the text question).
- If the answers are different: Great! The AI is actually looking at the picture.
- If the answers are the same: STOP. The AI is ignoring the picture. It doesn't matter if it's consistent or confident; it's not doing its job.
The Takeaway
In medical AI, consistency is not enough. An AI can be perfectly consistent and perfectly confident while being completely wrong because it isn't looking at the patient's scan.
To keep patients safe, we need to stop trusting AI just because it answers the same way every time. We need to make sure it's actually looking at the picture, not just guessing based on the words.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.