VLM-UQBench: A Benchmark for Modality-Specific and Cross-Modality Uncertainties in Vision Language Models
The paper introduces VLM-UQBench, a new benchmark designed to evaluate modality-specific and cross-modal uncertainty in vision-language models, revealing that current uncertainty quantification methods struggle to detect subtle, instance-level ambiguities and provide inconsistent signals for detecting hallucinations.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are hiring a personal assistant to help you navigate the world. You want them to be smart, but more importantly, you want them to be honest. If they can’t see a street sign clearly because it’s blurry, or if you ask a confusing question like, "What color is that thing over there?" without pointing, you don't want them to just guess and give you a wrong answer. You want them to say, "I'm not sure—is the image too blurry, or was your question too vague?"
In the world of Artificial Intelligence, these "assistants" are called Vision-Language Models (VLMs). They look at pictures and read text to answer questions. The problem? They are notorious "hallucinators"—they often confidently tell you there is a cat on the sofa when there is actually just a pillow.
This paper introduces VLM-UQBench, a high-tech "honesty test" designed to see if these AI models actually know when they are confused and, crucially, why they are confused.
The "Why" Problem: The Three Sources of Confusion
Most tests just ask, "Is the AI right or wrong?" This paper says that’s not enough. To fix an AI, we need to know where the mistake started. They categorize confusion into three buckets:
- The "Bad Eyes" Problem (Visual Uncertainty): The image is too dark, blurry, or pixelated. The AI is struggling because the "eyes" can't see.
- The "Bad Instructions" Problem (Textual Uncertainty): The human asked a confusing, typo-ridden, or overly subjective question. The AI is struggling because the "ears" didn't understand.
- The "Lost in Translation" Problem (Cross-Modal Uncertainty): The image is clear and the text is clear, but they don't match. For example, asking "What color is the car?" when there are three cars in the photo. The AI is struggling because it can't connect the words to the right object.
The "Stress Test" Pipeline
To test the AI, the researchers created a "digital torture chamber" (a perturbation pipeline). They take a perfectly clear photo and a perfect question, and then they systematically mess them up:
- They smudge the photo (Blur).
- They add typos to the text (Typos).
- They intentionally make the question and image clash (Cross-modal).
By doing this, they can see if the AI’s "uncertainty score" (its internal "I'm not sure" meter) actually goes up when things get messy.
The Surprising Results: The "Confident Liar"
The researchers tested several famous AI models (like GPT-4o-mini and LLaVA), and the results were a wake-up call. They found three main issues:
- Specialized Confusion: Some AI models are good at sensing when a question is bad, but totally blind to when an image is blurry. They are like a student who knows they didn't study the textbook but doesn't realize they can't read the handwriting on the exam.
- The Hallucination Gap: This is the most dangerous part. Even when the AI starts hallucinating (making things up), its "uncertainty meter" often stays low. It’s like a person walking confidently off a cliff because they are too certain the bridge is there.
- The "Subtle" Failure: AI models are okay at detecting "obvious" confusion (like a completely black image), but they are terrible at detecting "subtle" confusion (like a tiny typo or a slightly ambiguous reference).
Why This Matters
If we want to use AI in hospitals (to analyze X-rays) or in self-driving cars, "mostly right" isn't good enough. We need AI that can say, "Stop! I'm confused because the lighting is bad," or "Please rephrase that; I don't know which object you mean."
VLM-UQBench provides the roadmap for building that kind of reliable, self-aware AI. It moves us away from just building "smart" models and toward building trustworthy ones.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.