← Latest papers
💻 computer science

VidNum-1.4K: A Comprehensive Benchmark for Video-based Numerical Reasoning

This paper introduces VidNum-1.4K, a comprehensive human-annotated benchmark featuring 1,379 video-question pairs organized into a three-level hierarchy to rigorously evaluate and expose the significant limitations of current Vision-Language Models in performing complex, multi-step numerical reasoning within real-world video contexts.

Original authors: Shaoyang Cui, Lingbei Meng, Yaodi Luo, Peize He

Published 2026-07-20
📖 4 min read☕ Coffee break read

Original authors: Shaoyang Cui, Lingbei Meng, Yaodi Luo, Peize He

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are watching a movie, but instead of just enjoying the plot, you are being tested on your ability to count objects and events, calculate differences between quantities, and figure out if the number of people in a crowd grew or shrank after a scene cut. This is the world of Vision-Language Models (VLMs). Think of these models as super-smart digital detectives that can "see" images and videos and "read" text, trying to understand the world just like we do. For a while, these AI detectives have been getting really good at describing what they see, like saying, "There is a dog running." But scientists have been wondering: Do these AIs actually understand the story, the passage of time, and the logic of how things change? Or are they just guessing based on patterns they memorized from their training data? This question matters because if an AI can't truly understand the physical world—like knowing that if you add two apples to three, you have five, even if the camera angle changes—then we can't trust it to help us with complex tasks in the real world.

Enter a new challenge called VidNum-1.4K, a rigorous "final exam" designed specifically to test if these AI detectives can do numerical reasoning in videos. The researchers behind this study, Shaoyang Cui and their team, built a massive collection of 1,379 video clips and questions that force the AI to do more than just look; they must count, compare, and calculate. They organized these questions into three levels of difficulty, like a video game with increasing boss fights. Level 1 is the "easy mode," where the AI just has to count simple things, like "How many green doors are in this room?" Level 2 ramps up the difficulty, asking the AI to count specific types of things, like "How many brown dogs are there?" while ignoring the black ones, even if the camera cuts to different angles. Level 3 is the "hardcore mode," requiring the AI to perform multi-step logic, such as counting players in the first shot of a soccer game, counting them again in a second shot under specific constraints, and then performing a logical comparison or calculation between these two isolated events to find the answer.

The results of this exam were a bit of a shocker. Even the most advanced AI models, including the powerful closed-source "Gemini-3.1-pro," only managed to get about 60% of the answers right when allowed to think step-by-step. The open-source models, which are like the community-built tools of the AI world, struggled even more, scoring between 25% and 45%. The paper suggests that these models are still relying on "shortcuts"—like guessing based on common phrases they've heard before—rather than building a stable "internal world model" that understands how objects and numbers behave over time. For instance, when the video cuts to a new scene, the models often lose their count or double-count the same object, showing they don't truly grasp that the dog at the start of the clip is the same dog at the end.

Interestingly, the researchers found that asking the AI to "think out loud" (a technique called Chain-of-Thought) helped them solve the hardest logic puzzles, but it sometimes made them mess up the easier counting tasks they got right before. This suggests that while these models are getting better at high-level math, their basic ability to track moving objects in a video is still shaky. The study concludes that simply making the AI models bigger (adding more "parameters") helps them with complex logic, but it doesn't magically fix their trouble with tracking things across different video scenes. VidNum-1.4K serves as a tough, honest mirror, showing us that while our AI detectives are getting smarter, they still have a long way to go before they can truly understand the dynamic, changing world around them.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →