CaST-Bench: Benchmarking Causal Chain-Grounded Spatio-Temporal Reasoning for Video Question Answering
This paper introduces CaST-Bench, a new benchmark and evaluation suite designed to rigorously assess Vision-Language Models' ability to perform causal chain-grounded spatio-temporal reasoning in video question answering by providing a dataset of 2,066 questions with fine-grained temporal and spatial annotations.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are watching a mystery movie. Most video AI models today are like super-fast scanners that can tell you exactly what objects are on screen: "There is a woman. There is a child. There is a blue bag." They are great at describing what is happening.
But they are terrible at explaining why it is happening.
This paper introduces CaST-Bench, a new "final exam" designed to test if video AI can actually think like a detective, connecting the dots between cause and effect in real-time.
Here is a simple breakdown of what the paper does and finds:
1. The Problem: The AI is a "Smart Guessing Machine"
Current AI models often answer "Why" questions by guessing based on patterns they've seen before, rather than looking at the specific video evidence.
- The Analogy: Imagine a student taking a test. If you ask, "Why did the woman stop walking?", a smart-guessing AI might say, "Because she was tired," because that's a common reason people stop. But in the video, she actually stopped because her child fell behind. The AI missed the specific visual clue (the child falling) and relied on a "spurious correlation" (a lucky guess based on general knowledge).
2. The Solution: CaST-Bench (The "Causal Chain" Exam)
The authors built a new benchmark called CaST-Bench. Think of this as a rigorous training ground where AI models must prove they aren't just guessing.
- The Task: The AI is given a video and a "Why" question. To get a point, it can't just give an answer. It must build a Causal Chain.
- The Chain: This is like a trail of breadcrumbs. The AI must find:
- The Cause: (e.g., "The child squatted down at 00:05").
- The Effect: (e.g., "The woman stopped at 00:12").
- The Proof: It must point to the exact seconds and the exact spot on the screen (a bounding box) where these things happened.
If the AI says "The woman stopped to wait for her child," but it can't point to the video evidence of the child falling behind, it fails the test.
3. How They Built It: The Human-AI Team
Creating this exam was hard because videos are messy. The authors used a Human-AI Collaborative Pipeline:
- Step 1: They took real-world videos (not movies, but messy, real-life scenes) and used AI to track every person and object.
- Step 2: Humans reviewed the AI's descriptions to make sure they were accurate.
- Step 3: They used a "Causal Thinker" AI to generate tricky questions and answers.
- Step 4 (The Filter): They played a game of "Hide and Seek." They masked (blacked out) the parts of the video that contained the answer. If the AI could still answer the question correctly without seeing the key evidence, they threw that question away. This ensures the questions require the AI to look at the specific clues.
4. The Results: The AI is Still Stumped
The authors tested 15 different top-tier AI models (including big names like Gemini and GPT) on this new exam. The results were sobering:
- Humans: Scored 92%. We are good at connecting the dots.
- Best AI Models: Scored around 50%.
- The Gap: The models are struggling significantly. They often get the right answer but for the wrong reasons (hallucinating evidence), or they fail to point to the correct time and place in the video.
Key Finding: The paper shows that even the smartest AI models are bad at grounding. They can't reliably say, "I know this because I saw this specific pixel at this specific second."
5. Why This Matters (According to the Paper)
The paper argues that for AI to be truly useful and trustworthy, it needs to stop guessing and start proving.
- Transparency: If an AI can show you the exact video clip that led to its conclusion, you can trust it more.
- Accuracy: By forcing the AI to build a causal chain, it stops relying on lucky guesses and starts understanding the actual mechanics of the scene.
In a Nutshell:
CaST-Bench is a new "driver's test" for video AI. It doesn't just ask, "Can you see the car?" It asks, "Can you explain why the car stopped, and show me the exact moment the driver hit the brakes?" Currently, most AI drivers are failing this test, proving that we still have a long way to go before machines can truly understand the "why" behind what they see.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.