← Latest papers
💬 NLP

PaSBench-Video: A Streaming Video Benchmark for Proactive Safety Warning

The paper introduces PaSBench-Video, a streaming video benchmark demonstrating that current multimodal large language models fail to provide timely and accurate proactive safety warnings, as they struggle to distinguish emerging risks from safe scenes without triggering excessive false positives.

Original authors: Yusong Zhao, Yuejin Xie, Youliang Yuan, Junjie Hu, Jitian Guo, Yujiu Yang, Pinjia He

Published 2026-06-02
📖 5 min read🧠 Deep dive

Original authors: Yusong Zhao, Yuejin Xie, Youliang Yuan, Junjie Hu, Jitian Guo, Yujiu Yang, Pinjia He

Original paper dedicated to the public domain under CC0 1.0 (http://creativecommons.org/publicdomain/zero/1.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are the safety guard at a busy playground. Your job isn't just to yell "Stop!" when a child has already fallen off the slide. Your real job is to spot the wobble before the fall, shout a warning, and give the parent just enough time to catch the child.

This paper introduces a new "test" for AI cameras to see if they can do exactly that. The researchers call it PaSBench-Video.

Here is the breakdown of what they did, what they found, and why it matters, using simple analogies.

1. The Problem: The AI is a "Hindsight" Detective

Current AI safety systems are like detectives who only show up after the crime has happened. They look at a video and say, "Oh, a car crashed!" or "Oh, someone fell!"

But for safety to work, the AI needs to be a proactive alarm clock. It needs to see the danger building up (like a car braking hard or a toddler climbing a crib rail) and sound the alarm while there is still time to fix it.

The researchers found that existing tests for AI safety were flawed. They were like asking a driver to take a written test about a crash that happened yesterday, rather than watching them drive in real-time. The old tests didn't care when the AI yelled "Danger!"—only that it eventually did.

2. The New Test: A "Live Fire" Drill

The team created PaSBench-Video, a massive collection of 740 video clips. Think of this as a "driving simulator" for AI safety.

  • 481 clips show things going wrong (a car about to hit a cyclist, a person about to slip, a baby climbing a railing).
  • 259 clips show normal, safe scenes (cars driving smoothly, people walking in a park).

The Rules of the Test:

  1. Real-time only: The AI can only see what has happened so far. It cannot peek at the future.
  2. The "Goldilocks" Zone: The AI must yell "Warning!" in the perfect window of time.
    • If it yells too early (before the danger is visible), it's a false alarm (like a smoke detector going off because you burned toast).
    • If it yells too late (after the crash), it's useless.
    • It must also explain what is dangerous (e.g., "That car is braking," not just "Something is wrong").
  3. The Silent Test: If the video is safe, the AI must stay silent.

3. The Results: The AI is "Crying Wolf"

The researchers tested 13 of the smartest AI models available today. The results were sobering.

The "Wolf" Problem:
The AI models are terrible at timing. They are like a nervous guard who screams "Fire!" every time someone walks through a door.

  • The Trade-off: When the researchers told the AI to be more sensitive (to catch more real dangers), it started screaming at safe scenes constantly.
  • The Math: There is a tight link between catching real dangers and making false alarms. To catch 100% of the real dangers, the AI would have to scream "Danger!" on 60–80% of the safe videos. This would make the system useless because people would stop listening to it (a phenomenon known as "alarm fatigue").

The "Blind Spot" Problem:
The AI struggles most in driving scenarios.

  • Daily Life: In a home, danger often looks weird (a baby climbing a wall). The AI can spot this "weirdness" easily.
  • Driving: In traffic, a car merging safely looks almost identical to a car about to crash. The AI cannot tell the difference. It sees "cars moving" and panics, screaming warnings for normal traffic.

The "Wrong Reason" Problem:
Even when the AI does yell at the right time, it often gets the reason wrong.

  • Real Danger: "A brown car is pulling out."
  • AI Warning: "A blue bus is coming!"
    It sees the danger but hallucinates the details. It's like a security guard shouting "Intruder!" but pointing at the wrong person.

The Scorecard:
On the strictest test (correct timing + correct reason + no false alarms), no AI model scored higher than 20%. That means 8 out of 10 times, the AI either missed the danger, yelled too early, yelled too late, or yelled at the wrong thing.

4. The Reality Check: It's Not Ready for the Real World

The paper also looked at the practical side. Even if the AI were smart enough, it's not fast enough or cheap enough to be used right now.

  • Speed: To work in real-time, the AI needs to make a decision every 0.3 seconds. The fastest AI takes about 3.5 seconds to think. That's like a guard taking 10 seconds to react to a falling child.
  • Cost: Running this system on a single camera all day would cost thousands of dollars a day for the best models.

The Bottom Line

This paper is a "reality check" for AI safety. It shows that while AI is getting better at understanding videos, it is not yet smart enough to be a proactive safety guard.

Currently, these models are like a student who can describe a car crash perfectly after it happens, but cannot predict the crash before it occurs. They rely on surface-level "weirdness" rather than truly understanding how danger builds up. Until they can distinguish between a normal car and a crashing one without screaming at everything, they aren't ready to keep us safe on the road or in our homes.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →