Contextualized Evaluation of Vision Language Models through Dynamic, Multi-turn Interactions
The paper introduces CEDI, a dynamic, multi-turn interaction framework that bridges the gap between static benchmarks and real-world applications by revealing significantly more and more realistic visual hallucinations in Multimodal Large Language Models compared to conventional evaluation methods.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are trying to teach a robot how to see the world. For a long time, scientists have tested these robots by showing them a single picture and asking, "What do you see?" It's like a pop quiz where the robot has to write a short essay about a photo. If the robot gets the main things right, it gets an A. But in the real world, we don't just stare at photos and write essays. We talk to our devices. We ask, "Is that a dog?" then, "What color is its collar?" then, "Wait, is that a dog or a fox?" We have conversations. We change our minds. We get confused. The problem is that our old "pop quiz" tests don't catch the mistakes robots make when they are actually chatting with us. They might pass the quiz but fail the conversation, making up facts that aren't there just to keep the chat going. This is a big deal because if a robot is driving a car or helping a doctor, making things up can be dangerous. We need a way to test them the way we actually use them: in a messy, back-and-forth chat.
This paper introduces a new way to test these "Vision Language Models" (VLMs) called CEDI. Think of CEDI not as a test, but as a three-person play. First, you have the Robot (the model being tested). Second, you have a Detective (an automated examiner). Third, you have a Judge (a grader). The Detective doesn't just show the Robot a picture and walk away. Instead, the Detective is given a secret map of the picture (called a "Scene Graph") and a specific goal, like "Find the parking spot." The Detective then starts a conversation with the Robot, asking questions that get trickier and more specific as the chat goes on. The Detective might ask, "Is there a red car?" and if the Robot says yes, the Detective might follow up, "What model is it?" or even try to trick the Robot by asking, "Why is the blue car parked there?" when there is no blue car. The goal is to see if the Robot will start making things up (hallucinating) to keep the conversation alive.
The authors found that this "Detective" method is much better at catching lies than the old "pop quiz" style. When they used CEDI on several different robots, they discovered that the robots made up significantly more fake details than the old tests showed. For example, on one dataset, the new method found that the robots' "hallucination rates" jumped by 5 to 23 percentage points compared to the old methods. The robots were also much more likely to lie when the conversation got long. It seems that as the chat goes on, the robots get confused by their own previous answers and start reinforcing their own mistakes, like a game of "telephone" where the message gets more and more wrong.
The paper also showed that the robots are terrible at saying "I don't know." When the Detective asked a question based on a fake premise—like "What is the man holding?" when there is no man in the picture—the robots often tried to answer anyway, inventing a man and a hat. They struggled most when they were supposed to reject a false idea. The new "Judge" in this system, which uses a special graph-matching score, was able to catch these lies more consistently than previous tools, matching human judges' opinions better than the old methods did.
In short, the paper suggests that the old way of testing robots with single questions is missing a huge chunk of their problems. By switching to a dynamic, multi-turn conversation that mimics real life, we can see that these robots are much more prone to making things up, especially when they are trying to keep a conversation going or when they are tricked by false assumptions. The authors aren't saying the robots are broken, but they are saying we need to stop testing them like they are taking a silent exam and start testing them like they are having a conversation, because that's where the real trouble hides.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.