← Latest papers
💻 computer science

VideoZeroBench: Probing the Limits of Video MLLMs with Spatio-Temporal Evidence Verification

VideoZeroBench introduces a hierarchical benchmark with 500 manually annotated long-video questions and a five-level evaluation protocol to rigorously verify spatio-temporal evidence, revealing that even state-of-the-art video MLLMs like Gemini-3-Pro fail to achieve accurate grounded reasoning when precise temporal and spatial localization is required.

Original authors: Jiahao Meng, Tan Yue, Qi Xu, Haochen Wang, Zhongwei Ren, Weisong Liu, Yuhao Wang, Renrui Zhang, Yunhai Tong, Haodong Duan

Published 2026-04-03
📖 4 min read☕ Coffee break read

Original authors: Jiahao Meng, Tan Yue, Qi Xu, Haochen Wang, Zhongwei Ren, Weisong Liu, Yuhao Wang, Renrui Zhang, Yunhai Tong, Haodong Duan

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are hiring a detective to solve a mystery based on a 20-minute security camera tape.

The Old Way (Current Benchmarks):
You ask the detective, "Who stole the cookie?"
The detective says, "It was the butler!"
You check the answer key: Correct! You give them a gold star.
The Problem: You have no idea how they knew. Did they actually watch the tape? Did they see the butler's muddy footprints? Or did they just guess because butlers are often in movies? The current tests for AI video models are like this: they only check if the final answer is right, ignoring whether the AI actually "saw" the evidence.

The New Paper (VideoZeroBench):
The authors of this paper, "VideoZeroBench," decided to stop giving gold stars for lucky guesses. They built a much harder test to see if AI models are actually watching the video or just hallucinating (making things up).

Think of their test as a five-level detective exam:

  • Level 1 (The Cheat Sheet): You give the detective the video, plus a sticky note saying, "Look at minute 5:00 and minute 12:00, and look at the guy in the red hat."
    • Result: Even with these hints, the best AI models only get about 25% right. They still struggle to connect the dots.
  • Level 2 (The Time Hint): You take away the "red hat" hint, but you still say, "Look at minutes 5:00 and 12:00."
    • Result: The AI gets confused. It can't find the specific object in that time window.
  • Level 3 (The Standard Test): You just say, "Watch the video. Who stole the cookie?" (This is what all other AI tests do).
    • Result: The best AI (Gemini-3-Pro) gets about 17% right. That's barely better than a random guess!
  • Level 4 (The Time Proof): You demand, "Tell me the answer, AND show me the exact seconds on the tape where you saw it."
    • Result: The AI crashes. Accuracy drops to 8%. It can't point to the right time.
  • Level 5 (The Ultimate Test): You demand, "Tell me the answer, show me the exact seconds, AND draw a box around the specific object in the frame."
    • Result: Catastrophic failure. The best AI gets 1% right. Almost all models get 0%.

The Big Takeaway: The "Needle in a Haystack" Problem

The paper reveals a shocking truth: Current AI models are terrible at finding small details in long videos.

Imagine a 20-minute video is a giant haystack. The answer to the question is a tiny needle hidden somewhere in it.

  • Current AI: It looks at the haystack, guesses "It's probably a needle," and gets lucky sometimes. But if you ask it to find the needle, it fails.
  • The Specific Struggles:
    • Counting: If you ask, "How many red balls rolled by?" the AI often can't count them accurately.
    • Small Objects: If the clue is a tiny sign on a wall 10 minutes into the video, the AI misses it completely.
    • Direction: If a car turns left, the AI might think it turned right.

Why Does This Matter?

The authors argue that we are entering the "second half" of AI development. It's not enough for a model to be a "smart guesser." For AI to be truly useful (like a self-driving car or a medical diagnostic tool), it must be able to prove its reasoning. It needs to say, "I know this because I saw this specific frame at this specific time."

The Human Comparison

The researchers even tested humans on a small part of this exam.

  • Humans: Got 67% right.
  • Best AI: Got 22% right.

This huge gap shows that while AI is great at summarizing a movie or talking about the general plot, it is still very bad at being a precise, detail-oriented observer.

The Conclusion

VideoZeroBench is a wake-up call. It tells us that the "smart" video AIs we see in the news are actually quite "blind" when it comes to the fine details. To build truly reliable AI, researchers need to stop focusing just on getting the right answer and start teaching the models how to actually find and point to the evidence. Until they can do that, they are just very confident guessers.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →