← Latest papers
💬 NLP

Believing without Seeing: Quality Scores for Contextualizing Vision-Language Model Explanations

This paper proposes two new quality scoring functions, Visual Fidelity and Contrastiveness, to evaluate Vision-Language Model explanations, demonstrating that displaying these scores alongside explanations significantly improves users' ability to correctly assess prediction accuracy without visual context and reduces overreliance on incorrect outputs.

Original authors: Keyu He, Tejas Srinivasan, Brihi Joshi, Xiang Ren, Jesse Thomason, Swabha Swayamdipta

Published 2026-04-23
📖 4 min read☕ Coffee break read

Original authors: Keyu He, Tejas Srinivasan, Brihi Joshi, Xiang Ren, Jesse Thomason, Swabha Swayamdipta

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are blindfolded, and a friend (an AI) is describing a photo to you. They say, "This photo shows a sunny noon day." They sound very confident and give you a reason: "Look, the shadows are short, and there's a clock on the building showing 12:00."

Because you can't see the photo yourself, you have to decide: Do I trust my friend?

In the past, if your friend sounded convincing, you would believe them. But what if your friend is hallucinating? What if there is no clock, and the shadows actually suggest it's morning? You might blindly trust them and make a mistake.

This paper is about building a "Trust Meter" for AI assistants, especially for people who can't see the images the AI is looking at (like blind or low-vision users).

Here is the simple breakdown of their solution:

1. The Problem: "The Smooth-Talking Liar"

Current AI models are great at making up stories that sound true. Even when they get the answer wrong, they can generate a very logical-sounding explanation.

  • The Trap: If an AI says, "It's noon because the sun is high," and you can't see the sun, you might think, "Wow, that makes sense!" and trust the wrong answer.
  • The Old Way: Scientists tried to measure how "smart" or "informative" the explanation sounded. But they found that a smooth-talking liar can score high on these tests, too.

2. The Solution: Two New "Truth Detectors"

The authors propose two new ways to grade the AI's explanation, acting like a quality control inspector.

A. Visual Fidelity (The "Fact-Checker")

  • The Analogy: Imagine the AI is a tour guide describing a museum. Visual Fidelity asks: "Did the guide actually look at the paintings, or did they just make things up?"
  • How it works: The system breaks the AI's explanation down into small claims (e.g., "There is a clock," "The shadows are short"). It then uses a second AI to look at the photo and verify: Is there actually a clock? Are the shadows short?
  • The Score: If the AI mentions a clock that isn't there, the score drops. If it only mentions things that are actually in the photo, the score stays high.

B. Contrastiveness (The "Detective")

  • The Analogy: Imagine you are trying to find a specific person in a crowd. Contrastiveness asks: "Did the detective explain why this person is the one, and why everyone else is NOT the one?"
  • How it works: A good explanation doesn't just say "It's noon." It should explain why it isn't morning or afternoon. If the AI says "It's noon" but forgets to mention the clock (the key clue that proves it's noon and not morning), the score is low. It needs to show that it ruled out the other possibilities.
  • The Score: If the explanation clearly distinguishes the right answer from the wrong ones, the score is high.

3. The Experiment: Putting it to the Test

The researchers ran a study where human participants acted as the "blind" users. They were shown:

  1. A question about a photo they couldn't see.
  2. The AI's answer.
  3. The AI's explanation.
  4. Sometimes: A "Trust Score" (0 to 100) based on the two detectors above.

The Results were amazing:

  • When people saw the Trust Score, they got 11% better at spotting when the AI was lying.
  • Most importantly, the rate of people blindly believing a wrong answer dropped by 15%.
  • It was like giving the blindfolded user a "lie detector" test for their friend's story.

4. The Big Takeaway

You don't need to be a tech expert to know when an AI is hallucinating. By checking two simple things—"Is what you said actually in the picture?" and "Did you explain why the other options are wrong?"—we can give users a clear signal on when to trust the AI and when to be skeptical.

In short: This paper teaches us how to stop blindly believing the AI's smooth talk and start trusting the "receipts" (the facts) instead.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →