← Latest papers
💬 NLP

Visual Semantic Entropy: Do Vision Language Models Recognize Visual Ambiguity?

This paper proposes Visual Semantic Entropy (VSE), a novel uncertainty estimation method that isolates visual ambiguity by perturbing only image inputs while keeping text queries fixed, thereby overcoming the limitations of existing entropy-based approaches that are skewed by overconfident visual embeddings or dominated by textual variations.

Original authors: Ta Duc Huy, Trang Nguyen, Townim Chowdhury, Ankit Yadav, Minh-Son To, Zhibin Liao, Johan W. Verjans, Vu Minh Hieu Phan

Published 2026-07-01
📖 4 min read☕ Coffee break read

Original authors: Ta Duc Huy, Trang Nguyen, Townim Chowdhury, Ankit Yadav, Minh-Son To, Zhibin Liao, Johan W. Verjans, Vu Minh Hieu Phan

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are asking a very smart, confident robot to describe a picture. Sometimes, the picture is blurry or tricky, and the robot might be wrong. The big question this paper asks is: How can we tell if the robot is unsure, or if it's just confidently guessing?

The authors found that current ways of checking the robot's confidence are broken. They built a new, better way called Visual Semantic Entropy (VSE). Here is how they explain it using simple analogies:

The Problem: The "Overconfident Robot"

Think of a Vision-Language Model (VLM) as a robot that looks at a photo and answers a question.

  • The Old Way (Semantic Entropy): To check if the robot is unsure, we ask it the same question 10 times in a row. If it gives 10 different answers, we know it's confused. If it gives the same answer 10 times, we think it's sure.
  • The Flaw: The authors discovered that sometimes the robot is overconfident. Imagine the robot looks at a blurry photo of a bag that looks a bit like a pouch. Its "visual brain" is so convinced it's a pouch that even if you ask it 100 times, it will always say "pouch." It never wavers.
    • The Result: The old method sees the robot saying "pouch" 100 times and thinks, "Great, it's 100% sure!" But the robot is actually wrong and the picture was ambiguous. The old method missed the danger because the robot's internal confidence was too strong.

The Second Problem: The "Wordy Paraphrase" Trap

Some researchers tried to fix this by changing the question instead of just asking it repeatedly. They asked, "What is in the bag?" and then "What's inside the sack?" and "Describe the container."

  • The Flaw: The authors found that changing the words (text) causes the robot to get confused about the words, not the picture. It's like asking a friend, "Is that a cat?" and then "Is that a feline?" The friend might get confused by the different words and give different answers, even if the picture is perfectly clear.
    • The Result: The uncertainty score goes up, but it's measuring how sensitive the robot is to wording, not how ambiguous the image actually is.

The Solution: Visual Semantic Entropy (VSE)

The authors propose a new method that fixes both problems. Think of it as a "Visual Stress Test."

  1. Keep the Question, Shake the Picture: Instead of changing the words, they keep the question exactly the same. But, they slightly "jiggle" or "perturb" the image. Imagine taking a photo of a dog and adding a tiny bit of static noise or shifting the lighting just a hair.
    • Analogy: It's like looking at a blurry painting. If you step back, squint, or tilt your head (changing the view slightly), does the image still look like a dog, or does it start looking like a cat?
  2. Group the Answers by Meaning: The robot answers the same question for all these slightly different versions of the image.
    • If the robot says "pouch," "pocket," and "bag," a smart system realizes these all mean roughly the same thing. It groups them together.
    • If the robot says "pouch" for some views and "dog" for others, that's a real disagreement.
  3. Measure the Spread: They calculate how far apart these groups of answers are.
    • If the robot is truly unsure about the image, the groups will be far apart (e.g., one group says "pouch," another says "dog"). This creates a high uncertainty score.
    • If the robot is sure, all the groups will be close together, resulting in a low uncertainty score.

Why It Works Better

The paper tested this new method on five different smart robots and five different sets of tricky questions.

  • The Result: The new method (VSE) was much better at spotting when the robots were actually confused by a tricky image. It didn't get fooled by the robot's overconfidence, and it didn't get confused by changing the words.
  • The Takeaway: To know if a robot is unsure about a picture, you have to wiggle the picture, not the words. This gives a true measure of how ambiguous the visual evidence really is.

In short, the paper says: Don't trust a robot just because it repeats the same answer. If the picture is tricky, shake the picture a little bit. If the robot starts giving different answers, then you know it's unsure.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →