BRUCE: Benchmarking Robustness Under Corruption Escalation for Scientific Vision-Language Reasoning
This paper introduces BRUCE, a novel framework that benchmarks the robustness of scientific vision-language models against escalating image corruptions by employing new metrics to quantify performance degradation and providing an interpretable analysis of reasoning failures across OCR, spatial, symbolic, and semantic domains.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are teaching a super-smart robot to understand the world by showing it pictures and asking questions. This is the world of Vision-Language Models (VLMs). Think of these models as a student who has read every book in the library and can also see every photo ever taken. They are amazing at answering questions when the picture is crystal clear and the lighting is perfect. But in the real world, things aren't always perfect. Photos get blurry, the contrast might be low, or parts of the image might get cut off. The big question scientists are asking is: If you hand this super-smart student a slightly damaged or messy photo, will they still get the answer right, or will they completely fall apart?
This is where the idea of robustness comes in. It's not just about how smart the robot is when everything is easy; it's about how well it handles a little bit of chaos. Usually, when we test these robots, we only check if they get the answer right on perfect pictures. But this paper argues that's like only testing a car on a smooth racetrack and never seeing how it handles a pothole or a rainy day. We need to know if the robot's reasoning stays steady when the visual information starts to degrade, because in real life, images are rarely perfect.
The "BRUCE" Framework: Stress-Testing the Robot's Brain
In this paper, the researchers introduce a new testing system called BRUCE (Benchmarking Robustness Under Corruption Escalation). Imagine BRUCE as a very strict, slightly mischievous gym coach for these AI models. Instead of just letting the models answer questions, BRUCE starts messing with the pictures they are looking at. It doesn't just make one small change; it progressively makes the images worse and worse, step by step.
The coach might start by blurring the image a tiny bit, then a little more, then a lot. Or, it might start cropping the picture, slowly eating away the edges until only a tiny slice remains. It even tries "traversing" the image, which is like slowly covering the picture from top-to-bottom or left-to-right, removing information piece by piece to see exactly when the model starts to panic and give up.
To measure how well the models handle this stress, the authors created two new scorecards:
- RCI (Robustness Corruption Index): This measures how quickly the model's performance drops as the image gets messier. It's like a "fragility meter."
- T-RCI (Traversal-RCI): This is a special score for the "eating away" tests. It checks if the model collapses when information is removed from a specific direction, revealing if the model is relying on just one tiny part of the image to solve the whole problem.
What They Found: The "Smart" Student vs. The "Stable" Student
The researchers tested seven different AI models on tricky science and math problems, including chemistry diagrams and geometry puzzles. Here is what they discovered:
1. Being "Smart" Doesn't Mean Being "Sturdy"
The most surprising finding is that a model can be incredibly smart on perfect pictures but terrible when the picture is slightly damaged. For example, GPT-4.1 was great at solving chemistry problems when the images were clean, but as soon as the researchers started removing parts of the image (traversal), its performance crashed harder than almost any other model. It turned out GPT-4.1 was like a student who memorized the exact layout of a diagram but couldn't figure out the logic if a piece was missing.
2. The "Specialist" Isn't Always the Best
They tested ChemVLM-8B, a model specifically trained for chemistry. You might think a chemistry specialist would be the best at chemistry problems. However, the results showed it wasn't necessarily more robust than general-purpose models. In fact, it struggled just as much, if not more, when the visual evidence was corrupted. This suggests that just knowing chemistry facts isn't enough; the model also needs to be good at "seeing" through the noise.
3. The "Stable" Winners
Two models stood out as the most resilient: InternVL2.5-8B and Qwen2.5-VL-72B. These models didn't just get the right answers; they kept getting them even when the images were blurry or had large chunks missing. Qwen2.5-VL-72B, in particular, had the highest clean accuracy (87.23%) and maintained a high accuracy (84.18%) even when the images were corrupted. It seems these models have a more distributed way of thinking, where they don't rely on just one part of the image to solve the puzzle.
4. The "Direction" Matters
The study found that it matters how you remove information. For some models, removing the image from the top-down was fine, but removing it from the right-to-left caused them to fail immediately. This suggests that these AI models have "blind spots" or biases in how they look at a picture. They might be looking at the right side of the image for clues, and if you cover that side, they are lost.
5. The Real Problem: Reasoning, Not Just Seeing
When the images got worse, the models didn't just stop "seeing" the objects. Instead, they stopped "reasoning." The researchers found that the biggest type of failure was Symbolic Reasoning. This means the models could still see the shapes and numbers, but they couldn't connect the dots to solve the math or science problem. It's like a student who can read the numbers on a clock but can't figure out what time it is if the hands are slightly blurry.
The Takeaway
The paper concludes that we can't just look at how well an AI does on perfect tests. We need to stress-test them with messy, broken, and incomplete images to see if their reasoning holds up. The authors suggest that for AI to be truly useful in the real world—where photos are often blurry or cut off—we need models that are not just smart, but also robust enough to handle the messiness of reality without losing their train of thought. The "BRUCE" framework gives us the tools to find out which models are truly ready for the real world and which ones are just good at taking perfect test photos.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.