Diagnosing Corruption-Induced Reliability Failures in Vision-Language Models
This paper introduces Bench-C, a controlled testbed and the Robustness Alignment Score (RAS) metric to reveal how visual corruptions can silently degrade the distributional reliability and structural stability of Vision-Language Models, even when top-1 accuracy appears to improve.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you're taking a multiple-choice quiz with a super-smart robot friend. You show it a picture of a red apple and ask, "What color is this?" The robot confidently shouts, "Red!" and gets a gold star. But what if the picture was blurry, or covered in static, or looked like it was taken on a rainy day?
Most people would still say, "It's definitely red!" even if the photo was a bit messy. But here's the twist: Vision-Language Models (VLMs)—the robots that look at pictures and talk about them—can get tricked in ways that a simple "Right or Wrong" score completely misses.
That's exactly what this paper, titled "Diagnosing Corruption-Induced Reliability Failures in Vision-Language Models," is all about. The authors built a special testing ground called BENCH-C to peek under the hood and see what happens when the robot's vision gets "corrupted" (like when a photo gets blurry, noisy, or distorted).
The Big Surprise: The Robot Can Lie to Itself
The main finding is a bit counterintuitive. Usually, we think if a robot gets a question wrong, it's broken. But this paper found that sometimes, the robot gets the answer "right" even when its internal confidence is totally falling apart.
Think of it like this:
- The Silent Crash: The robot sees a blurry picture of a red apple. It still says "Red!" (so the score says "Correct!"). But inside, its brain is screaming, "I'm not sure! It could be red, or orange, or maybe a tomato!" It lost its certainty, but the final answer didn't change. The score says "Good," but the reliability is actually broken.
- The Lucky Guess: Sometimes, a messy picture makes the robot change its mind from "Blue" to "Red" (the correct answer). The score says "Great! It got better!" But the paper suggests this might just be a fluke. The robot didn't suddenly understand the apple; it just stumbled into the right answer while its brain was wobbling.
The authors argue that just looking at the final answer (Top-1 Accuracy) is like judging a movie only by the ending. You might miss the fact that the actors were forgetting their lines, the camera was shaking, or the plot made no sense until the very last second.
The New Tool: BENCH-C and the "RAS" Score
To catch these sneaky failures, the team built BENCH-C. Imagine a giant bag of 4,000 quiz questions. Most of them are boring because the robot gets them right no matter how bad the picture is (like asking "Is this a cat?" when the picture is clearly a cat).
The authors filtered this bag down to 849 "diagnostic" questions. These are the tricky ones where the robot's answer actually changes or gets shaky when the picture gets messy. They tested these on 13 different VLMs using 19 different types of visual corruption (like blur, noise, weather, and digital glitches) at 5 levels of severity (from "a little bit fuzzy" to "unrecognizable").
To measure the robot's "mental state," they invented a new score called RAS (Robustness Alignment Score).
- If the robot gets the answer right and stays confident, RAS is happy.
- If the robot gets the answer right but is confused (or gets it wrong but is super confident), RAS drops.
What They Found (The "Aha!" Moments)
The paper measured 13 models and found some weird patterns:
Mild corruption can actually make the score look better, even if the robot is worse.
Sometimes, adding a little bit of noise to a picture makes the robot's answer change from "Wrong" to "Right." The score goes up! But the authors suggest this isn't real improvement; it's just unstable behavior. The robot is wobbling, and sometimes it wobbles into the right answer by accident.The "Source" of the problem matters.
They split the questions into two groups: ones the robot got right before the picture got messy, and ones it got wrong before.- Originally Correct: When these got messy, the robot often kept the right answer but lost its confidence (Silent Degradation).
- Originally Wrong: When these got messy, the robot sometimes fixed its answer, but often it was just a shaky, temporary fix.
Severity changes the game.
As the corruption got worse (from level 1 to 5), the damage to the "Originally Correct" questions got much worse. But strangely, the "Originally Wrong" questions sometimes looked like they were getting better (more likely to guess right). The authors suggest this isn't because the robot is learning; it's because the chaos is shaking things up so much that it occasionally hits the right answer by luck, while the solid answers are crumbling.
What They Didn't Find (The "Nope" List)
The paper is very careful about what it doesn't claim:
- It's not about how often these errors happen in real life. The testbed (BENCH-C) was designed to find sensitive samples, not to count how many real-world photos are blurry. So, don't assume this means 50% of all photos are broken; it just means these specific photos are great for testing.
- It's not a magic fix. The paper doesn't say they solved the problem or that the robots are now safe. It just says, "Hey, our current way of testing is missing a lot of the danger."
- It's not about the robot's brain being "conscious." They are looking at math (probability distributions), not feelings.
The Takeaway
The authors measured 80,655 corrupted image-question pairs and found that reliability is more than just getting the right answer. A model can look perfect on a report card while its internal logic is falling apart.
They suggest that in the future, we shouldn't just train robots to get the right answer. We should train them to stay confident when they are right and stay unsure when they are wrong, even when the picture is a mess.
In short: Don't just trust the answer. Trust the confidence behind it. If the robot says "Red!" but its brain is shaking, it's not a reliable friend yet.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.