Chains That See, Answers That Don't: A Multi-Aspect Evaluation Recipe for Forced Chain-of-Thought on Video-MME
This paper introduces a multi-probe evaluation framework that reveals while forced chain-of-thought reasoning in video-language models is genuinely conditioned on visual input, it fails to improve—and may even slightly degrade—multiple-choice question accuracy on the Video-MME benchmark.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you have a very smart, but sometimes over-eager, robot assistant. You show it a video and ask a question about it. To make the robot seem more trustworthy, you tell it: "Before you give me the answer, please write down your step-by-step thinking process first." This is called "Chain-of-Thought" (CoT).
The big assumption in the AI world is that this forced thinking process makes the robot smarter and more accurate. It's like believing that if a student writes out their math work, they are less likely to make a mistake than if they just guessed the answer.
This paper is like a detective story where the authors put that assumption to the test. They didn't just ask, "Did the robot get the right answer?" They asked, "Is the robot actually thinking about the video, or is it just reciting a script?"
Here is the story of their investigation, broken down into three simple experiments:
The Setup: The "Video-MME" Test
The authors used a giant library of video clips and multiple-choice questions (like a TV quiz show). They tested two versions of the robot: a "Big Brain" (32 billion parameters) and a "Small Brain" (7 billion parameters).
Experiment 1: The "Script Check" (Did the robot actually watch?)
The Analogy: Imagine you ask a student to write an essay about a movie they just watched.
- The "Boilerplate" Fear: What if the student just memorized a generic essay template? They write, "The movie had great acting and a good plot," regardless of whether they watched Titanic or The Matrix. They haven't actually seen the movie, but they still wrote a long essay.
- The Test: The authors showed the robot a video, then swapped it for a completely different video (e.g., swapping a baseball game for a cooking show) but kept the same question.
- The Result: The robot did not use a generic script. When the video changed, the robot's "thinking" changed completely.
- On the real video, it said: "I see two Spider-Men drinking tea."
- On the swapped video (a baseball game), it said: "This video is about baseball, not Spider-Men. I cannot answer."
- The Takeaway: The robot was watching the video. Its "thinking" was real and specific to what it saw. It wasn't faking it.
Experiment 2: The "Thinking vs. Doing" Test (Did the thinking help?)
The Analogy: Now that we know the student is actually watching the movie, does writing down the essay help them get the right answer on the quiz?
- The Test: They compared two groups:
- Group A: "Just tell me the answer."
- Group B: "Write your thinking steps, then tell me the answer."
- The Result: Surprisingly, writing the thinking steps did not help. In fact, for the "Small Brain" robot, it made things worse.
- The "Big Brain" robot got about the same score either way.
- The "Small Brain" robot got significantly fewer correct answers when forced to write the steps.
- The Takeaway: Even though the robot was "thinking" correctly about the video, that extra step of writing it down actually confused the smaller robot. It's like a student who knows the answer but gets so nervous trying to write out their logic that they mess up the final calculation. The "thinking" was real, but it didn't translate into a better score.
Experiment 3: The "Blurry Vision" Test (How much visual info is used?)
The Analogy: Imagine you are taking a test while wearing glasses that get progressively dirtier.
- The Test: They showed the robot videos that were:
- Clear and real.
- Shuffled (frames in random order).
- Single frame (just one picture repeated).
- Black screen (no picture at all).
- The Result: As the video got "dirtier" or less informative, the robot's "thinking" text changed to match the loss of information, and its accuracy dropped.
- The Takeaway: This confirmed again that the robot was truly looking at the visual input. If it were just reciting a script, a black screen wouldn't change its "thinking" text.
The Big Conclusion
The paper's main message is a bit of a twist: Just because a robot writes a long, logical explanation that proves it "saw" the video, doesn't mean that explanation helps it get the right answer.
- The Chains That See: The robot's thinking process was deeply connected to the video (it wasn't a fake script).
- The Answers That Don't: Despite seeing the video and writing the steps, the forced thinking process did not improve the final score. For the smaller model, it actually hurt performance.
In short: The authors built a special "recipe" to test AI. They found that forcing an AI to "show its work" proves it is looking at the picture, but it doesn't guarantee it will get a better grade. Sometimes, just asking for the answer directly is actually more reliable than forcing a long, step-by-step explanation.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.