← Latest papers
🤖 AI

See More, Think Deeper: Query-Expanded Visual Evidence and Answer-Clue Guided Reflection for Long Video Understanding

The paper introduces CoVER, a framework that enhances long-video understanding in Video-LLMs by dynamically gathering query-expanded visual evidence to "see more" and employing answer-clue guided reflection to "think deeper," thereby outperforming existing models through evidence-centric and visually verifiable reasoning.

Original authors: Shuning Wang, Zhiheng Wu, YiNuo Lu, Naiming Liu, Chen Jia, Bowen Liu, Shuo Nie, Weijie Zhu, Yumeng Zhang

Published 2026-06-09
📖 4 min read☕ Coffee break read

Original authors: Shuning Wang, Zhiheng Wu, YiNuo Lu, Naiming Liu, Chen Jia, Bowen Liu, Shuo Nie, Weijie Zhu, Yumeng Zhang

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to solve a mystery in a very long, complex movie. You are the detective, and the movie is your only source of clues.

Most current "AI detectives" (Video-LLMs) try to solve the mystery by watching the movie once, very quickly, and then guessing the answer. They might miss a tiny detail because they were looking at the whole screen at once, or they might guess based on what usually happens in movies rather than what actually happened in this specific scene.

The paper introduces a new detective named CoVER (Comprehensive Visual Evidence and Reflection). CoVER doesn't just guess; it follows a strict two-step process to "See More" and "Think Deeper."

Here is how CoVER works, using a simple analogy:

The Problem: The "Glance and Guess" Approach

Imagine you are asked, "What was the very first snack the character ate?"

  • Old AI: Watches the whole movie quickly. It sees a character eating an apple later on. It guesses, "It was an apple!" It didn't look closely enough to see that the character actually drank a soda with ice before eating the apple.
  • The Issue: The AI relies on a single, broad look and often misses the small, crucial clues hidden in the timeline.

The Solution: CoVER's Two-Step Detective Work

CoVER changes the game by using two special tools: The Magnifying Glass and The Fact-Checker.

Step 1: "See More" (The Magnifying Glass)

Instead of just watching the movie once, CoVER acts like a detective who realizes, "I need to look closer."

  • The Trick: Before answering, CoVER asks itself a bunch of follow-up questions, which the paper calls "pseudo-queries."
  • The Analogy: If the question is "What did they eat first?", CoVER doesn't just look for "food." It generates specific search terms like: "Is there a drink with ice?", "Are there watermelon slices?", "Is there a cup?"
  • The Result: It uses these specific questions to zoom in on tiny, high-definition clips from the video that it might have missed during the first quick glance. It gathers a complete pile of evidence before making a guess.

Step 2: "Think Deeper" (The Fact-Checker)

Once CoVER has gathered the evidence, it makes a Draft Answer (a first guess). But it doesn't stop there.

  • The Trick: CoVER takes its own draft answer and turns it into a "Clue" to test itself.
  • The Analogy: If CoVER guesses, "The first snack was watermelon," it creates a specific clue: "Check if the watermelon appears before the soda." It then goes back to the video specifically to look for that relationship.
  • The Result: If the video shows the soda came before the watermelon, CoVER catches its own mistake. It says, "Oh no, my draft answer was wrong because the evidence contradicts it," and it revises the answer to the correct one (the soda).

Why This Matters

Most AI models are like students who write an essay and hand it in immediately. CoVER is like a student who:

  1. Researches the topic deeply (gathering more evidence).
  2. Writes a first draft.
  3. Critically reviews their own draft against the facts.
  4. Fixes mistakes before handing in the final paper.

The Results

The paper shows that this method works. CoVER is much better at answering questions about long videos than other models of the same size. It even beats some very expensive, "closed-source" models (models you can't see inside) on certain tests.

In short: CoVER stops AI from guessing based on a quick glance. Instead, it forces the AI to hunt for specific clues, write a draft, and then double-check that draft against the video to ensure the answer is actually true.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →