← Latest papers
🤖 AI

Visual Credit Audit for Multimodal Spatial Reasoning

This paper introduces Visual Credit Audit (VCA), a training- and label-free framework that decomposes multimodal spatial reasoning performance into correctness, visual support, and relation-specific response, revealing that a significant portion of correct benchmark answers are actually uncredited by the visual evidence.

Original authors: Feixiang Liu, Qiang Qiu, Lanbo Sun, Nan Wei, Huawei Shen, Xueqi Cheng

Published 2026-07-30
📖 6 min read🧠 Deep dive

Original authors: Feixiang Liu, Qiang Qiu, Lanbo Sun, Nan Wei, Huawei Shen, Xueqi Cheng

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are taking a test where a teacher shows you a picture of a cat sitting under a TV and asks, "Is the TV above the cat?" You answer "Yes," and you get a gold star. But what if you got that gold star not because you actually looked at the picture, but because you guessed right based on the words alone? Maybe you've heard that phrase a million times, or maybe you just know that TVs usually go on walls and cats on floors. In the world of Artificial Intelligence, specifically with "Multimodal Large Language Models" (AI that can see and read), this is a tricky problem. These models are like super-smart students who have read every book in the library, but sometimes they cheat by guessing the answer from the question's wording rather than actually analyzing the image. Scientists want to know: When the AI says "Yes," is it really seeing the TV above the cat, or is it just a lucky guess based on text?

This is exactly the puzzle tackled in a new paper called "Visual Credit Assignment for Multimodal Spatial Reasoning." The researchers built a special audit tool called Visual Credit Audit (VCA). Think of VCA as a strict detective who doesn't just check if the answer is right, but investigates why it was right. They ask two main questions: First, did the picture actually give the AI extra help to make that decision compared to just reading the text? Second, if you swapped the relationship in the picture (like putting the cat above the TV), would the AI change its mind? The paper finds that even when AI models get the right answer, they are often "cheating" by relying on text clues rather than visual evidence. In fact, for some models, up to 26.25% of their correct answers on spatial tests were actually uncredited—they got the right answer, but the image didn't actually help them get it.

The Detective's Toolkit: How VCA Works

To understand how the researchers caught these "cheaters," imagine a game of "Spot the Difference" played with a very stubborn robot.

The Setup: The "No-Image" Controls
The researchers set up a clever experiment. They took a question like "Is the TV above the cat?" and showed it to the AI in three different ways:

  1. The Original: The question plus the real picture.
  2. The Text-Only: The question with no picture at all (just a blank space).
  3. The Blank Canvas: The question with a gray, empty box where the picture should be.

If the AI answers "Yes" in all three scenarios with the same confidence, the researchers say, "Aha! The picture didn't do any work." The AI is just guessing based on the text. But if the AI is much more confident when the picture is there, then the picture gets "credit" for the answer. This is the first part of the audit: Image Dependence.

The Twist: The "Relation Reversal"
But wait, there's a second trap. What if the AI is sensitive to the picture, but only in a weird way? Maybe it sees the TV and the cat, but it doesn't actually understand which one is on top. To catch this, the researchers used a "factorial" test. They created scenarios where the text said "Is the TV above the cat?" but the picture showed the TV below the cat.

If the AI is truly smart, it should say "No" when the picture contradicts the text. If it still says "Yes" because it's ignoring the picture, it fails the second test: Relation Consistency.

The Big Reveal: "Correct but Uncredited"

The researchers ran this audit on four different AI models (Qwen, InternVL, LLaVA, and Ministral) using two different spatial reasoning benchmarks. The results were eye-opening.

They found that a significant chunk of the AI's "correct" answers were actually Correct-but-Uncredited (C-U).

  • On the GSR-COCO dataset, the Qwen model got 90.91% of the answers right. But when they applied the audit, only 78.18% of those answers were actually supported by the image. That means 12.73% of the time, the model got the right answer, but the image didn't help it at all.
  • For the LLaVA model on the same dataset, the gap was even wider: 26.25% of its correct answers were uncredited.

The paper explicitly rules out the idea that these models are just "bad" at spatial reasoning. In fact, the models are sensitive to visual changes. When the researchers swapped the positions of objects in the pictures (a "relation reversal"), 81.57% to 100.00% of the models did change their answers in the correct direction. This proves the models can see the difference, but in the standard test, they often didn't need to look to get the right answer. They were relying on shortcuts.

Why This Matters (Without the Jargon)

Think of it like a student taking a math test.

  • Accuracy is just the grade: "Did they get the right number?"
  • VCA is the teacher looking at the student's scratch paper: "Did they actually do the math, or did they just guess the answer because they knew the teacher likes the number 42?"

The paper shows that current AI benchmarks are like a test where the teacher accidentally gives away the answers in the question phrasing. The models are getting high scores, but they aren't necessarily learning to "see" better.

The researchers also tested what happens if you swap the picture with a completely different, unrelated image (like swapping the cat/TV photo with a photo of a beach). When they did this, the models' "credit" score dropped by 21.25 to 47.80 points. This confirms that the original picture was indeed doing the heavy lifting for those specific correct answers, and without it, the models would have stumbled.

The Takeaway

The paper doesn't say AI is broken or that it can't see. Instead, it offers a new way to measure success. It suggests that we shouldn't just count how many times an AI gets the right answer. We need to count how many times the AI needed the picture to get that answer.

By separating "correctness" from "visual support," the researchers found that a lot of what we think is "smart vision" is actually just "smart guessing." They propose that future tests for AI should include these "credit audits" to ensure that when an AI says "I see the cat," it really means it's looking at the cat, and not just reciting a phrase it memorized. It's a call to stop celebrating the grade and start checking the homework.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →