Hazard or Anomaly? Evaluating VLMs for Understanding Dangers and Discrepancies
This paper introduces a critical distinction between "hazard" and "anomaly" to evaluate Vision-Language Models, revealing that current models often misinterpret scene irregularities as dangers and demonstrating that separating these concepts provides a more accurate assessment of safety reasoning than traditional binary judgments.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are the captain of a high-tech rescue robot, tasked with scanning a chaotic disaster zone to keep people safe. Your robot's brain is a "Vision-Language Model" (VLM), a super-smart AI that can look at a picture and describe what it sees, kind of like a very chatty, observant friend. But here's the tricky part: in a real emergency, your robot needs to know the difference between something that is just weird and something that is deadly. A pile of colorful, abandoned toys might look strange and out of place, but it won't hurt anyone. A pile of live wires, however, is dangerous, even if it looks perfectly normal. If your robot can't tell the difference between "odd" and "dangerous," it might scream "FIRE!" at a red costume, distracting rescuers from the real threat, or worse, ignore a hidden trap because it looks too familiar. This paper dives into whether our current AI robots are smart enough to make that crucial distinction, or if they are just guessing based on how "unusual" a scene looks.
The researchers behind this study, Murali Indukuri and his team, decided to put these AI brains to the test with a new kind of exam. Instead of just asking the robots, "Is this safe or unsafe?" (a simple yes-or-no question), they asked them to sort scenes into four specific buckets: Safe (everything is normal), Anomalous (something is weird or out of place, but safe), Hazardous (something is dangerous, but looks normal), and Anomalous-Hazardous (something is both weird and dangerous). They created a special dataset of 610 images to act as the test questions.
When they ran the tests, they found a funny but worrying habit in the robots. The AI models were like over-enthusiastic alarmists who confused "strange" with "scary." If a scene looked unusual—like a toaster sitting in a bathtub—the AI often shouted "DANGER!" even if the toaster was unplugged and perfectly safe. The study suggests that these models are relying too much on the feeling that a scene is "irregular" to decide if it's dangerous, rather than actually understanding the physics of the situation. It's as if the robot thinks, "That doesn't look right, so it must be a monster!"
The team also tried different ways of asking the questions to see if they could fix this. They used three main strategies:
- Zero-Shot: Just giving the robot the definitions and asking it to guess.
- Few-Shot: Showing the robot a few examples of "weird but safe" and "dangerous" scenes before asking it to judge a new one.
- Chain-of-Thought: Asking the robot to talk through its reasoning step-by-step, like solving a math problem out loud.
The results showed that while the robots got better at spotting actual dangers (hazards) with the right prompts, they still struggled to separate the "weird" from the "deadly." The best-performing models, like GPT-4.1 and Gemini 3 Flash, managed to get about 85% accuracy on spotting hazards, but they were much less confident when identifying things that were just strange but safe. Even the smartest models made mistakes, sometimes missing a real danger or flagging a harmless oddity as a threat.
Interestingly, the researchers also tested if they could just give the robots a written description of the picture (a "dense caption") instead of the picture itself. Surprisingly, the robots did worse with just the text! It seems that for this specific job, seeing the actual image is still crucial, and reading a description isn't enough to help the AI understand the context.
The paper concludes that while these AI models are getting pretty good at their jobs, they aren't ready to be the sole guardians of safety-critical situations yet. They tend to overreact to anything that looks out of the ordinary. The authors suggest that by forcing the AI to explicitly separate "weird" from "dangerous" in its training and testing, we can help it become more reliable. They also found that their new, more detailed way of testing was much better at catching these mistakes than the old "safe vs. unsafe" tests. While the current models are a step in the right direction, the study suggests we need to keep refining how we teach them to think, ensuring they don't panic at a clown's red nose while ignoring a real fire.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.