← Latest papers
🤖 AI

When Seeing Is Not Enough: Benchmarking Interactive Visual Grounding in LVLMs

This paper introduces a benchmark for interactive visual grounding in large vision-language models (LVLMs), revealing that current models significantly underperform human baselines, struggle with proactive information seeking when no initial description is provided, and exhibit poor calibration between confidence and accuracy.

Original authors: Zhengxiang Wang, Owen Rambow

Published 2026-08-26
📖 5 min read🧠 Deep dive

Original authors: Zhengxiang Wang, Owen Rambow

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

In the world of artificial intelligence, there is a specific skill called visual grounding. It is the ability for a computer to look at a picture and point to the exact object a human is talking about. If you say, "Find the red ball," the system must locate that specific ball among many others. For years, scientists have tested this skill by giving the computer a perfect, detailed description and asking it to find the match. This is like handing someone a complete map and asking them to find a hidden treasure. The computer has become very good at this one-step task, often matching or beating human performance when the instructions are clear and complete.

However, real life is rarely that simple. When humans talk to each other, descriptions are often vague, incomplete, or even wrong. We do not hand each other perfect maps; instead, we build our understanding together. We ask questions like, "Is it on the left?" or "Does it have a handle?" to fill in the missing pieces. This paper explores what happens when we ask artificial intelligence to do the same thing: to look at a picture, realize it does not have enough information, and then ask the right questions to find the target. The researchers wanted to see if these smart computer systems could handle the messy, back-and-forth nature of real conversation, or if they would fail when the instructions were not handed to them on a silver platter.

The team set up a controlled experiment to test this interactive ability. They created a game with two roles: a "Director" who knows the target object and an "Matcher" who must find it. They used four different visual worlds for the game: a collection of dogs, baskets, and abstract shapes called tangrams. In the game, the Matcher sees a grid of many possible items, while the Director sees only the correct target. The researchers tested the Matcher under four different conditions to see how it handled information. In the first condition, the Matcher was given a full, perfect description and had to find the object without talking. In the second, it was given a full description but allowed to ask follow-up questions to check its work. In the third, it started with a very vague clue, like "it is a basket," and had to ask questions to figure out which one. In the final, most difficult condition, the Matcher was given no description at all and had to ask yes-or-no questions from scratch to build a picture of the target in its mind.

They ran this game with eight different large vision-language models, which are the most advanced types of artificial intelligence that can see and read. The results were clear and somewhat surprising. Even when the models were given perfect descriptions, they struggled to match the performance of human players. But the gap grew much wider when the task required interaction. When the models had to ask questions to find the target, their performance dropped significantly. They were particularly poor at the hardest task, where they had to start with no information and guide the conversation themselves. While they could ask questions, they often asked the wrong ones or failed to use the answers they received to narrow down their choices effectively.

The study also looked at how confident the models were in their answers. The researchers found that the models were often overconfident. They would report a high level of certainty that they had found the right object, even when they were actually wrong. This lack of self-awareness was most pronounced when the models had to ask questions to find the target. They seemed to believe they had enough information when they did not. Interestingly, the models did show some ability to improve when they were allowed to ask follow-up questions to refine a description, but this help was limited. They could fix a mistake if someone pointed it out, but they were not good at proactively seeking the information they needed to avoid the mistake in the first place.

To ensure these findings were solid, the researchers ran several follow-up tests. They checked if the results changed if the descriptions were written by a computer instead of a human, if the models were forced to think harder before answering, or if they played the game multiple times with the same partner. In every case, the pattern remained the same. The models performed better with computer-generated descriptions, but they still struggled with the interactive parts of the task. They also found that while the models could get slightly better at the game after playing it a few times, they never caught up to human performance. The difficulty of starting a conversation from scratch and building a shared understanding remained a major hurdle. One follow-up study did test the models on face images, but the main game focused on dogs, baskets, and tangrams.

The researchers concluded that while these artificial intelligence systems are excellent at matching a clear description to a picture, they are not yet ready for the real world of human communication. They lack the ability to recognize when they are missing information and to ask the right questions to get it. They are also poor at judging how sure they should be about their answers. This suggests that for robots or AI assistants to truly work alongside humans, they need to learn more than just how to see; they need to learn how to talk, how to ask for help, and how to realize when they do not know the answer. Until they can do that, they will remain limited to tasks where the instructions are perfect, leaving the complex, interactive nature of human reference as a challenge they have yet to solve.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →