← Latest papers
🤖 AI

Accuracy Without Grounding: Diagnosing Visual Dependency Dissociation in Video LLM Benchmarks

This paper introduces the Visual Dependency Gap (VDG) to demonstrate that high benchmark accuracy in video large language models often stems from language priors rather than genuine visual understanding, revealing that temporal order contributes minimally to performance while frame diversity drives most gains.

Original authors: Jae Joong Lee

Published 2026-07-16
📖 5 min read🧠 Deep dive

Original authors: Jae Joong Lee

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to teach a robot to understand the world. You give it a camera and a brain, and you ask it questions about what it sees. For a long time, scientists have measured how "smart" these robots are by giving them a test: show them a video, ask a question, and see if they get the right answer. If the robot gets a high score, we assume it is truly "seeing" and understanding the video. But here is the catch: just because a robot gets the right answer doesn't mean it looked at the picture. It might have just guessed based on the words in the question, like a student who knows the teacher always asks about "cats" on Tuesdays and just answers "cat" without looking at the blackboard. This paper dives into the messy, confusing world of "Video Large Language Models"—super-smart computers that can read text and watch videos—to ask a scary question: Are these robots actually watching the movies, or are they just really good at guessing the answers using only the text?

The researchers decided to play a little trick on twenty different video-watching robots, ranging from tiny ones to massive ones with huge brains. They took a standard test called MVBench, which has hundreds of questions about videos, and they ran the test twice for every single question. The first time, they showed the robot the real video. The second time, they showed the robot the exact same question, but the video was replaced with a solid black screen. If the robot was truly "seeing," its score should drop to near zero when the video disappears. If the robot kept getting the answers right even with a black screen, it meant it wasn't watching at all; it was just using language tricks to guess.

The results were a bit of a shock. The team found that for many questions, the robots performed almost exactly the same whether they were watching a video or staring at a black void. They introduced a new way to measure this called the "Visual Dependency Gap" (VDG), which is basically the difference between how well a robot does with a video versus without one. They discovered that for some types of questions, like "What color is the car?" (Attribute Perception), the robots really did need the video to get it right. But for other types, like "What happens next?" (Temporal Reasoning), the robots got the answers right even when the screen was black. This means the test questions weren't actually testing if the robots understood time or sequence; they were just testing if the robots could guess the answer based on the wording of the question.

The researchers also tried to figure out why the robots were getting these answers. They broke the video down into different parts: just one single frame, a bunch of frames shuffled out of order, and the full video with time moving forward. They found that the robots didn't care about the order of events at all. Whether the video played forward, backward, or was just a jumbled mess of pictures, the robots got the same score. The only thing that seemed to help was seeing different pictures (frame diversity), not seeing them in the right order. In fact, for some of the newest, most advanced models, the video input actually seemed to confuse them, making them perform slightly worse than if they had just ignored the video entirely and guessed based on the text.

One of the most interesting findings was about how these robots handle compression. Usually, if you squish a video file to make it smaller, the quality gets worse. The researchers expected that if a robot was truly "seeing," a blurry, compressed video would make it fail. But the robots' scores stayed flat and steady even when the video was heavily compressed. The team realized this wasn't because the robots were super-robust; it was because the questions were so easy to guess from the text that the robots didn't need the visual details at all. The "robustness" was an illusion created by bad test questions.

Finally, they looked at how these robots changed as they got bigger and smarter. They found that while bigger models generally got better at using the video, some of the very newest models actually got worse at relying on visual information. One model, Qwen3-VL, seemed to have forgotten how to use the video frames effectively, relying almost entirely on guessing from the text, even though its overall test scores looked good. The paper concludes that we can't just trust the leaderboard scores anymore. A high score doesn't mean a robot is a great observer; it might just mean it's a great guesser. To know if a robot is truly seeing, we need to check if it fails when the screen goes black. If it doesn't fail, it wasn't watching in the first place.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →