EgoCoT-Bench: Benchmarking Grounded and Verifiable Operation-Centric Chain of Thought Reasoning for MLLMs
The paper introduces EgoCoT-Bench, a fine-grained egocentric video benchmark featuring 3,172 verifiable QA pairs with explicit step-by-step rationale annotations, designed to evaluate and improve the grounded, operation-centric reasoning capabilities of Multimodal Large Language Models.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are watching a video of someone cooking from their own point of view (like wearing a GoPro on their forehead). You see hands grabbing a knife, a pot, and a spoon. Now, imagine you ask a super-smart AI robot: "Where is the salt shaker after the chef puts it down?"
The robot might say, "It's on the counter." And it might be right! But here's the catch: Did the robot actually see the salt shaker move there, or did it just guess?
This is the problem the paper EgoCoT-Bench is trying to solve.
The Problem: The "Lucky Guess" Robot
Right now, many AI models are like students who memorize the answer key but don't understand the math. They can look at a video and give the correct answer to a question, but their explanation (their "reasoning") might be made up, inconsistent, or based on things that didn't actually happen in the video.
In the world of "First-Person" (Egocentric) videos, this is extra hard because:
- Hands often block the view.
- Objects disappear and reappear.
- The camera shakes and moves wildly.
Existing tests for AI just check if the final answer is right. This paper says, "That's not enough! We need to check if the AI's story about what happened is actually true to the video."
The Solution: EgoCoT-Bench (The "Truth Detective" Test)
The authors built a new, super-detailed test called EgoCoT-Bench. Think of it as a "driving test" for AI, but instead of driving a car, the AI has to understand a video of someone doing tasks with their hands.
How they built it:
- The Map (STSG): Instead of just watching the video, they first built a "map" of the video. This map tracks every object, every hand, and every second of time, like a detailed script of a play.
- The Questions: They used this map to create 3,172 questions. These aren't just "What color is the cup?" questions. They are tricky ones like: "The hand grabbed the blue cup at 1:00, then the red cup at 1:05. Which one was on the table at 1:03?"
- The Proof: Every single question comes with a "receipt." The test includes the exact time, the location, and the visual proof needed to answer it. This way, we can check if the AI's reasoning matches the receipt.
The 4 Types of Challenges
The test is divided into four main categories, like levels in a video game:
- Spotting the Action (Grounding & Perception): "What is the hand touching right now?" (Can the AI see what's happening now?)
- The Time Machine (Retrospection): "Where was the spoon before the hand picked it up?" (Can the AI remember the past?)
- The Fortune Teller (Prediction & Causal Inference): "The hand is holding the lid. What will happen next?" or "Why did the coffee spill?" (Can the AI guess the future or explain the cause?)
- The Big Picture (High-Level Reasoning): "Is the chef done making the sandwich, or is there one more step?" (Can the AI understand the whole goal?)
What Happened When They Tested the AI?
The authors took the smartest AI models available (like GPT-5, Qwen, and LLaVA) and put them through this test. Here is what they found:
- The "Lucky Guess" Phenomenon: Many models got the right answer (e.g., "The bottle is on the shelf"), but when asked why, their explanation was wrong. They might have said, "Because I saw a bottle," when the video actually showed the bottle being moved after the question time.
- The Gap: Even the best AI models scored around 60-70% correct. Humans scored nearly 96%. This means there is still a huge gap between human understanding and AI understanding when it comes to watching hands move things around.
- The Hardest Part: The AI was okay at guessing what happens next, but terrible at tracking an object's history (like following a specific tag on a shirt as it gets pulled off).
The Bottom Line
The paper introduces a new way to grade AI. Instead of just giving a grade based on the final answer, they now grade the reasoning too.
They found that many AIs are "spurious correct"—they get the right answer for the wrong reasons. This is dangerous because if you ask an AI to help a robot in a kitchen, and the AI guesses the right answer but has the wrong idea of where the knife is, the robot might cut something it shouldn't.
EgoCoT-Bench is a tool to force AI to stop guessing and start paying attention to the actual evidence in the video, step-by-step.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.