← Latest papers
💬 NLP

Seeing Is Not Sharing: Some Vision-Language Models Overestimate Common Ground in Asymmetric Dialogue

This paper reveals that vision-language models tend to overestimate mutual understanding in collaborative dialogue by conflating the mere presence of shared task-relevant content with established common ground, a bias driven by static referential cues rather than dynamic dialogue history.

Original authors: Nan Li, Albert Gatt, Massimo Poesio

Published 2026-07-01
📖 5 min read🧠 Deep dive

Original authors: Nan Li, Albert Gatt, Massimo Poesio

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

The Big Idea: "Seeing" Isn't the Same as "Sharing"

Imagine you and a friend are trying to draw a route on two different maps. You both have a map, but they aren't identical. Maybe your map has a "Parked Van" at the top, while your friend's map has one at the bottom.

When you say, "Go to the parked van," you are thinking of the one at the top. Your friend is thinking of the one at the bottom. You haven't actually agreed on which van yet, even though you used the same words.

In human conversation, we fix this by talking back and forth: "Wait, do you mean the one near the park?" "Oh, no, I mean the one by the bakery." This process of checking and confirming is called grounding.

The Problem:
The researchers asked: Can modern AI models (Vision-Language Models) tell the difference between "what we could share" and "what we have shared"?

They tested this by acting as a "third-party observer" (an overhearer). The AI could see the maps and hear the conversation, but it couldn't ask questions or interrupt to clarify.

The Main Discovery: The AI is Too Optimistic

The researchers found that when they gave the AI the actual maps, it got better at guessing the answer overall, but it became dangerously overconfident.

The Analogy: The "Co-Presence" Trap
Imagine you are watching a play from the balcony. You see two actors standing next to a red chair.

  • Actor A thinks the chair is for the villain.
  • Actor B thinks the chair is for the hero.
  • They haven't spoken about the chair yet.

Because you (the observer) can see the chair is right there for both of them, your brain assumes they must be talking about the same thing. You think, "Oh, they are definitely on the same page!"

The AI does the exact same thing. When it sees a landmark (like a "Parked Van") on both maps, it assumes the speakers have already agreed on it. It confuses potential (the object exists on both maps) with reality (the speakers have actually confirmed they mean the same one).

The Experiments: What Changed the AI's Mind?

The researchers tested the AI under different conditions to see what caused this "over-optimism."

1. The Visual Trap (Maps vs. Text)

  • Real Maps: When the AI saw the actual pictures of the maps, it said "Yes, they agree!" way too often.
  • Text Descriptions: When they replaced the pictures with a simple text list saying "Map A has a van; Map B has a van," the AI still said "Yes, they agree!" too often.
  • Blank/Shuffled Maps: When they showed the AI blank gray squares or scrambled map pieces with no useful info, the AI became very cautious. It stopped guessing "Yes" and started guessing "No."

The Lesson: The problem isn't that the AI is bad at "seeing" pictures. The problem is that it is bad at interpreting information. Whether the info comes as a picture or a text list, if the AI sees that an object could be shared, it assumes it is shared.

2. The "Overhearer" Effect
The AI is structurally an "overhearer." It hears every word but can't participate.

  • Humans know that just because two people are looking at the same object doesn't mean they are talking about it together.
  • The AI thinks: "I see the object on both maps, so they must be talking about it together."

It treats the existence of a shared object as proof of mutual understanding.

The Results: A Trade-Off

The AI's behavior created a specific pattern:

  • When the speakers actually agreed: The AI got very good at spotting this (especially with maps).
  • When the speakers were confused or hadn't agreed yet: The AI got terrible at spotting this. It confidently said "They agree!" even when they were actually talking past each other.

It's like a student who is great at answering questions they already know the answer to, but when they are confused, they confidently guess the wrong answer because they think the question is obvious.

Which AI Models Did This?

The researchers tested several models.

  • Qwen3-VL-8B: This model showed the bias the most clearly. It was very confident in its wrong answers when maps were involved.
  • Gemma3 Models: These behaved differently. Some were so confused by the complex map drawings that they barely tried to answer at all.
  • Size Doesn't Matter: Making the model bigger didn't fix the problem. In fact, the bigger models were often more confident in their wrong answers.

The Bottom Line

The paper concludes that current AI models are excellent at static analysis (looking at a map and seeing what could be shared) but poor at dynamic tracking (watching a conversation to see what has been shared).

They mistake the possibility of agreement for actual agreement. They see the "common ground" on the map and assume the speakers have already built a bridge across it, even if the speakers haven't said a word about it yet.

In short: The AI sees the map, assumes everyone is on the same page, and confidently ignores the fact that the speakers might be reading different chapters entirely.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →