Do Video-LLMs Actually Watch? Diagnosing Character-Tracking Failures in Long-Form Video
This paper reveals that current Video-LLMs largely fail to genuinely track specific characters in long-form videos, instead relying on shallow gender cues and failing to distinguish between same-gender individuals, which causes benchmark scores to overestimate their actual temporal reasoning and identity-tracking capabilities.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you've hired a super-smart, 7-to-8-billion-parameter robot detective to watch a whole TV episode of The Big Bang Theory. Your mission for the robot is simple: "Watch Sheldon, and tell me the exact order his outfits change."
You'd expect the robot to be like a hawk, locking onto Sheldon's face, tracking him through every scene, and noting his wardrobe changes with perfect precision. But here's the twist: the robot isn't actually watching Sheldon. It's just guessing based on a very lazy shortcut.
The Great "Sheldon" Swap
To figure out what was going on, the researchers played a clever trick. They took the exact same video and the exact same multiple-choice answers, but they simply swapped the name in the question.
- Original Question: "In what order does Sheldon change outfits?"
- The Swap: "In what order does Leonard change outfits?" (Same video, same answers).
If the robot was truly tracking characters, it should have said, "Wait, Leonard wears different clothes than Sheldon!" and picked a different answer. But guess what? The robot didn't care.
In about 7% to 17% of same-gender swaps (like swapping Sheldon for Leonard), the robot changed its answer. That's barely better than rolling a die. It was mostly ignoring the name and just picking an answer it liked.
The "Gender" Glitch
So, if it's not looking at the person, what is it looking at? The researchers found the robot was using a very coarse, blurry filter: Gender.
When they swapped a male character for a female character (like Sheldon to Penny), the robot was much more likely to change its answer (20% to 43% of the time). It seemed to think, "Oh, the question changed from a 'guy' to a 'girl,' so the clothes must be different!"
But if you swapped one guy for another guy, the robot was completely lost. It couldn't tell Sheldon from Howard any better than it could tell a cat from a dog. It was seeing "male" and "female" but couldn't see "Sheldon."
The "Magic Box" Illusion
The paper also checked if the robot was cheating by memorizing the show from its training data (like having read the script beforehand). They tested this by showing the robot only the text of the question and the show's title, with no video at all. The robot scored basically at random chance (around 18–20%). So, it wasn't cheating with a script; it was just bad at the visual part.
They also tried giving the robot more frames (more "snapshots" of the video) to see if it just needed to see more. Giving it 32 frames instead of 16 made it slightly better at spotting clothes in general, but it still didn't get better at tracking the specific character. It was like giving a blind person a better pair of glasses; they can see the clothes better, but they still don't know who is wearing them.
The Multiple-Choice Trap
Here's the most surprising part: The robot was getting 37–38% of the answers right on the standard test. That sounds impressive! But that score was a mirage.
When the researchers asked the robot to write the answer out in its own words (Open-Ended) instead of picking from a list (Multiple Choice), the score plummeted by 18 to 25 points.
- Multiple Choice Score: ~38%
- Open-Ended Score: ~13–19%
Why the drop? Because the multiple-choice test gave the robot a "magic box" of five options. Even if the robot was just guessing or picking the answer it liked most (like "E" or "A"), it had a 1-in-5 chance of being right. When forced to generate the answer from scratch, it had no safety net, and it failed completely. In fact, zero of the 151 open-ended answers from the open-source robots were fully correct.
The Bottom Line
The paper concludes that these Video-LLMs are not actually tracking characters. They are "hallucinating" a connection between the name and the video. They are good at spotting that "a person is wearing a shirt," but they are terrible at answering "Which person?"
The researchers suggest that the current high scores on benchmarks are misleading. They are measuring how well the models can guess based on gender cues or lucky multiple-choice guesses, not how well they can actually follow a story. Until a model can pass the "Name Swap" test and the "Open-Ended" test, we can't say it truly "watches" the video.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.