KnowHal: A Knowledge-Driven Benchmark for Comprehensive Multimodal Hallucination Evaluation
The paper introduces KnowHal, a comprehensive benchmark that unifies the evaluation of entity, attribute, relation, and knowledge hallucinations in Multimodal Large Language Models through a novel paired-question design, revealing that knowledge-related errors and false-premise acceptance remain significant challenges for current models.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are teaching a super-smart robot to understand the world. This robot, called a Multimodal Large Language Model (or MLLM for short), is like a brilliant student who can look at a picture and read a book at the same time. It's amazing at describing what it sees, like saying, "That's a red apple on a wooden table." But sometimes, this robot gets a little too creative. It might look at a picture of a cat and confidently tell you, "This cat is flying a spaceship," even though cats can't fly and there's no spaceship in the photo. This is called a "hallucination." It's like the robot is daydreaming while it's supposed to be working.
Why does this matter? Because we want to trust these robots with important jobs, like helping doctors diagnose illnesses or guiding self-driving cars. If a robot starts making things up, the results could be dangerous. Scientists have been trying to build tests to catch these daydreams. Most of these tests check if the robot can correctly identify simple things in a picture, like "Is there a dog?" or "Is the dog brown?" But there's a tricky part: sometimes the robot knows the facts about the world (like "dogs have tails") but gets confused when the picture doesn't match, or it invents facts that aren't in the picture at all. Until now, we didn't have a single, fair test that checked both what the robot sees and what it knows about the world, all in one go.
Enter KnowHal, a new and clever test designed to catch these daydreams in a whole new way. Think of KnowHal as a "spot the difference" game, but for robot brains. The researchers built a massive collection of 1,800 picture-and-question pairs, covering 10 different areas of life, from sports to music. For every single picture, they created two types of questions: "Positive" questions that ask about things actually in the photo or true facts about the object, and "Negative" questions that are tricky traps. These traps sound very plausible but are actually lies. For example, if a picture shows a real dog, a positive question asks, "What color is the dog?" A negative trap might ask, "What color is the dog's invisible blue hat?"
The researchers tested 14 different robot brains on this new game. They found some surprising things. First, the robots were generally okay at answering the "Positive" questions, but they stumbled badly when faced with the "Negative" traps. It's like a student who knows their math facts perfectly but gets tricked by a question that asks, "If 2 plus 2 equals 5, what is 2 plus 2?" The robots often said "5" instead of correcting the lie. This shows that even the smartest robots aren't very good at saying, "Wait, that's not true," when they are being misled.
The biggest surprise was that the hardest part for almost every robot was the "Knowledge" dimension. This is where the robot has to use facts it learned from books or the internet, not just what it sees in the picture. Even the best robots got this wrong more often than they got simple visual questions wrong. It suggests that while these robots are getting better at seeing, they are still struggling to separate what they see from what they think they know.
The study also looked at how the size of the robot brain affected its performance. They found that bigger brains generally did better, but even the biggest ones still struggled with the tricky "Knowledge" questions and the "Negative" traps. This tells us that just making the robot bigger isn't a magic fix; we need to teach them how to be more careful and skeptical.
In short, the paper introduces KnowHal, a new benchmark that pairs real questions with tricky, false ones to test how well robots can tell truth from fiction. The results suggest that while our AI friends are getting smarter, they still have a long way to go before they can be fully trusted to spot a lie or admit when they don't know the answer. The authors believe this new test will help future robots become more reliable and less prone to making up stories.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.