← Latest papers
💻 computer science

Know-Show: Benchmarking Video-Language Models on Spatio-Temporal Grounded Reasoning

This paper introduces Know-Show, a new benchmark and a training-free method called GRAM, to evaluate and improve the spatio-temporal grounded reasoning capabilities of Video-Language Models by unifying reasoning with visual and temporal localization across diverse scenarios.

Original authors: Chinthani Sugandhika, Chen Li, Deepu Rajan, Basura Fernando

Published 2026-04-01
📖 5 min read🧠 Deep dive

Original authors: Chinthani Sugandhika, Chen Li, Deepu Rajan, Basura Fernando

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are watching a home security video. A person walks into the kitchen, opens the fridge, takes out a carton of milk, and pours it into a glass.

If you ask a human, "Who poured the milk and when did they do it?", they can instantly point to the person, say "It was the guy in the blue shirt," and pinpoint the exact moment on the clock. They know the answer and can show you the proof.

Now, imagine asking a super-smart AI robot the same question. The paper you shared, titled "Know-Show," reveals a funny but serious problem: The robot might say, "A person poured milk!" (The Know part is okay). But if you ask, "Show me exactly who and when," the robot might point to the wrong person, guess the wrong time, or just stare blankly. It knows the story, but it can't ground that story in the actual video evidence.

Here is a simple breakdown of what the researchers did to fix this.

1. The Problem: The "Hallucinating" Detective

Current AI video models are like detectives who read a mystery novel but have never seen the crime scene. They are great at guessing the plot ("The butler did it!") because they've read millions of stories. But when you ask them to point to the specific moment the butler picked up the gun, they get lost.

  • The Gap: They can recognize objects (a fridge, a hand) and understand time (first this, then that), but they struggle to link them together. They can't say, "The left hand of Person A touched the fridge handle at 12:05 PM."
  • The Result: In the real world (like self-driving cars or medical monitoring), this is dangerous. If a car AI thinks a pedestrian is crossing but can't point to where or when, it might crash.

2. The Solution: The "Know-Show" Benchmark

The researchers created a new test called Know-Show. Think of this as a final exam for video AI, but with a twist.

Instead of just asking, "What happened?", the exam asks:

  1. Know: What is happening?
  2. Show: Point to the exact person, object, or hand involved, and tell me the exact second it happened.

The test has five levels of difficulty, like a video game:

  • Level 1: Point to the person doing the action.
  • Level 2: Point to the object being used.
  • Level 3: Point to both the person and the object together.
  • Level 4 (The Boss Level): Point to the specific hand touching the object. (This is very hard because hands move fast and are small).
  • Level 5: Tell me the exact order of events and when they started.

The Score: The AI only gets points if it gets the answer and points to the right spot at the right time. If it gets the answer right but points to the wrong person, it gets zero.

3. The Results: AI vs. Humans

When they ran this test on the smartest AI models available (like GPT-4o, Gemini, and others), the results were humbling:

  • Humans: Scored over 75%. We naturally connect actions to specific people and times.
  • AI Models: Scored very low (often under 20-30%). They were great at guessing the general idea but terrible at the "proof." They often confused who was doing what, or missed the exact moment an action started.

4. The Fix: GRAM (The "Flashlight" Plugin)

The researchers didn't want to retrain the massive AI models (which takes years and millions of dollars). Instead, they built a "plug-in" called GRAM.

Think of GRAM as a flashlight and a time-stamp marker for the AI:

  • The Flashlight (Attention Selection): When the AI starts to reason ("Okay, who opened the fridge?"), GRAM forces the AI to look at the specific video frames where the fridge is actually being opened. It filters out the blurry background and focuses only on the relevant pixels.
  • The Time-Stamp Marker: The AI usually guesses time based on a vague feeling. GRAM literally writes the time (e.g., "5.2 seconds") into the AI's thought process, forcing it to align its answer with the clock on the video.

The Outcome: When they added this "flashlight" to the AI, its performance jumped significantly. It didn't become perfect, but it finally started to "show what it knows" instead of just guessing.

The Big Picture

This paper is a wake-up call. It tells us that while AI is getting better at "talking" about videos, it still struggles to "see" and "prove" what it's talking about.

  • Before: AI was like a storyteller who makes up details.
  • After (with Know-Show & GRAM): We are teaching AI to be a forensic analyst who must back up every claim with visual evidence and a timestamp.

This is crucial for the future. If we want robots to help in hospitals, drive cars, or monitor safety, they can't just guess; they need to know exactly who did what and when, and be able to show us the proof.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →