← Latest papers
💻 computer science

From Accuracy to Visual Dependence: Auditing and Filtering Modality Collapse in Traffic VideoQA

This paper audits traffic accident VideoQA benchmarks to reveal that high accuracy often stems from textual shortcuts rather than visual evidence, proposing new metrics and a training-free filtering method to identify and remove shortcut-prone questions, thereby ensuring genuine visual grounding in safety-critical applications.

Original authors: Sena Korkut, María Alejandra Bravo Sarmiento, Sanghwan Kim, Zeynep Akata

Published 2026-06-30
📖 4 min read☕ Coffee break read

Original authors: Sena Korkut, María Alejandra Bravo Sarmiento, Sanghwan Kim, Zeynep Akata

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are taking a driving test. The examiner shows you a video of a traffic accident and asks, "What caused this crash?"

Ideally, you should watch the video carefully, see the car swerve, notice the pedestrian stepping out, and then answer based on what you saw. This is what we want AI to do: look at the evidence.

However, this paper reveals a problem: some AI models are "cheating." They aren't actually watching the video. Instead, they are reading the question and the multiple-choice answers, guessing the right answer based on patterns in the words alone. It's like a student who memorizes the answer key without ever studying the textbook.

Here is a breakdown of what the researchers found and how they fixed it, using simple analogies.

1. The Problem: The "Blind" AI

The researchers tested several AI models on four different traffic accident video quizzes. They discovered a strange phenomenon they call "Modality Collapse."

  • The Analogy: Imagine a detective trying to solve a crime. If the detective closes their eyes, puts on noise-canceling headphones, and still solves the case perfectly, you know they didn't actually look at the crime scene. They just guessed based on the description of the crime.
  • The Finding: On one specific test called MM-AU, the AI models actually got better scores when the video was removed! When they were forced to look at the video, their scores went down. This means the video was actually confusing them, and they were relying entirely on "text shortcuts" (guessing based on word patterns).

2. The Diagnosis: Two New Rulers

To measure how much an AI is actually "looking" versus just "guessing," the authors invented two new rulers (metrics):

  • The "Blind Gap" (How good are you at guessing?): This measures how well the AI does when it is blindfolded (text only). If the AI gets a high score while blindfolded, it means the questions have "shortcuts" that don't require looking.
    • Analogy: If a student can pass a math test just by reading the multiple-choice options without doing the math, the test has a high "Blind Gap."
  • The "Visual Gain" (How much does the video help?): This measures how much the AI's score improves when it is allowed to see the video.
    • Analogy: If giving the student a calculator (the video) makes their score go up, the calculator is useful. If the score goes down because the calculator is distracting, the test is broken.

The Results:

  • MM-AU: Had a huge "Blind Gap" and a negative "Visual Gain." The AI was cheating, and the video was useless.
  • TrafficQA: Had a low "Blind Gap" and a high "Visual Gain." The AI actually needed the video to get the right answer.

3. The Solution: The "Shortcut Score"

The researchers didn't just want to measure the problem; they wanted to fix the test. They created a "Shortcut Score" for every single question.

  • How it works: They gave the AI a question in two ways: once blindfolded and once with the video.
    • If the AI got it right while blindfolded with high confidence, the question gets a high Shortcut Score (it's a bad question).
    • If the AI only got it right after seeing the video, the question gets a low Shortcut Score (it's a good, visual question).
  • The Filter: They used this score to throw away the "cheating" questions and keep only the ones that truly required looking at the video.

4. The Result: A Fairer Test

After filtering out the bad questions:

  • The AI could no longer cheat.
  • The "Blind Gap" disappeared (the AI couldn't guess the answers without the video).
  • The "Visual Gain" became positive (the video actually helped the AI).

The Big Takeaway:
The paper argues that high accuracy is a trap. Just because an AI gets 90% of the answers right doesn't mean it understands the video. It might just be a master of word-guessing. For safety-critical tasks like traffic accidents, we need to ensure the AI is actually looking at the scene, not just reading the question.

The authors also noted that the reason some tests were so easy to cheat on was that the questions and answers were too repetitive. It's like if every multiple-choice question had the same answer format; the AI would learn the pattern, not the content. By cleaning up the questions, they forced the AI to actually "see."

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →