LongVQUBench: Benchmarking Long-Term Video Quality Understanding of Vision-Language Models
This paper introduces LongVQUBench, a comprehensive benchmark comprising over 1200 diverse long-duration videos and 1500 questions designed to evaluate the long-term video quality understanding of vision-language models through three hierarchical reasoning levels, revealing significant performance limitations in current state-of-the-art models regarding temporal integration and perceptual attribution.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you have a very smart robot that can watch videos and talk about what it sees. You ask it, "Did that video look blurry?" or "Why did the picture flicker?" For short clips, like a 10-second cat video, this robot is usually pretty good. But what happens when you ask it to watch a full two-hour movie and tell you if the quality got worse over time?
According to the paper LongVQUBench, current "super-smart" video robots are actually quite bad at this long-term quality check. Here is a simple breakdown of what the researchers did and found.
The Problem: The "Short Memory" Robot
Think of existing video robots like a person who has a very short attention span. They are great at noticing a smudge on a window if you point right at it (a short clip). But if you ask them to watch a whole day of surveillance footage and tell you when the camera started to get foggy, they get lost. They forget the beginning by the time they reach the end, and they struggle to connect small glitches into a big picture of "bad quality."
The Solution: A New "Stress Test"
The researchers built a new test called LongVQUBench. Think of this as a giant, tricky exam designed specifically to see how well these robots handle long videos.
- The Content: They gathered over 1,200 videos. Some are movies, some are news, some are shaky home videos, and some are cartoons.
- The Trick: They didn't just show the videos; they secretly planted "bugs" (distortions) in them. Imagine a video where, for a few seconds, the colors shift weirdly, or the picture stutters, or it gets grainy.
- The Needle in the Haystack: They call this the "Needle Distortion" test. It's like hiding a tiny needle inside a giant haystack. The robot has to find that tiny glitch in a video that might be two hours long.
The Three Levels of the Exam
The test gets harder in three steps, like climbing a ladder:
Level 1: The "Spot the Glitch" Test (Local Event)
- The Task: "Look at this 20-second slice. Is there a blur here?"
- Analogy: It's like asking a student, "Is there a typo in this one sentence?" Most robots can do this okay.
Level 2: The "Connect the Dots" Test (Cross-Event)
- The Task: "There was a glitch at minute 5 and another at minute 45. How do they compare? Did they make the whole video worse together?"
- Analogy: Now you ask the student, "Compare the typo in sentence 1 with the typo in sentence 50. Which one was worse for the story?" The robots start to struggle here because they have to remember the first part while looking at the second.
Level 3: The "Overall Vibe" Test (Global Understanding)
- The Task: "After watching the whole two-hour movie, was the quality good, bad, or did it get worse over time? Why?"
- Analogy: This is like asking the student to write an essay on the entire book, explaining how the writing quality changed from the first chapter to the last. This is where the robots fail the hardest. They can't synthesize the whole experience into a single, smart opinion.
What They Found
The researchers tested 14 of the smartest video robots available (including big names like GPT-5 and various open-source models).
- The Longer, The Worse: As the videos got longer and the questions got more complex, the robots' scores dropped significantly.
- More Frames ≠ Better: You might think, "If I show the robot more pictures from the video, it will do better." The researchers found that simply feeding the robot more frames didn't fix the problem. The robots still got confused about the timeline.
- The "Global" Gap: The robots were decent at spotting a single glitch (Level 1) but terrible at judging the overall quality of a long video (Level 3). They are like a person who can see a crack in a brick but can't tell you if the whole wall is about to fall down.
The Bottom Line
The paper concludes that while these AI models are amazing at understanding what is happening in a video (the story, the actions), they are still very clumsy at understanding how well the video looks over a long period. They lack the ability to keep a "mental map" of quality changes over hours.
The researchers created LongVQUBench to be a standard ruler for measuring this specific weakness, hoping it will help engineers build robots that can truly appreciate (or critique) a full-length movie without losing their train of thought.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.