LVLMs and Humans Ground Differently in Referential Communication
This paper presents a referential communication experiment demonstrating that Large Vision-Language Models (LVLMs) struggle to establish common ground and interactively generate or resolve referring expressions as effectively as humans, thereby limiting their ability to collaborate seamlessly with human users.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you and a friend are playing a game where you have to match 12 specific baskets from a pile of 16. You can't see what your friend sees, and they can't see what you see. You have to describe the baskets to each other to get them in the right order.
This is exactly what researchers did in this new study, but they swapped their human friends for AI robots (specifically, advanced Large Vision Language Models, or LVLMs). They wanted to see if AI can learn to "speak the same language" as humans or as other AIs over time, just like real people do.
Here is the breakdown of what they found, using some simple analogies.
The "Secret Handshake" of Conversation
When two humans play this game, something magical happens.
- Round 1: You might say, "Pass me the tall, brown basket with the handle that looks like a question mark."
- Round 2: You say, "The question-mark handle one."
- Round 3: You just say, "The Q-basket."
Humans naturally develop a secret handshake (called a "conceptual pact"). They realize, "Oh, we both know what 'Q-basket' means now, so I don't need to describe the whole thing again." They get faster, shorter, and more accurate. They are building a shared understanding, or common ground.
The AI Experiment: 4 Different Teams
The researchers set up four teams to play this game for four rounds:
- Human + Human: Two people.
- AI + AI: Two robots talking to each other.
- Human + AI: A person talking to a robot.
- AI + Human: A robot talking to a person.
The Results: Who Won?
1. The Humans (The Champions) 🏆
The human teams got better and better. By the end, they were speaking in shorthand, making fewer mistakes, and finishing quickly. They were like a jazz band that finally gets into a groove; they stopped over-explaining and started improvising together.
2. The AI vs. AI Team (The Verbose Robots) 🤖🤖
This was the most surprising result. Even when two AIs talked to each other, they did not get better.
- The Problem: They kept describing the baskets in excruciating detail every single time.
- Round 1: "The tall, brown, woven basket with a curved handle..."
- Round 4: "The tall, brown, woven basket with a curved handle..." (They said the exact same long sentence again!)
- The Analogy: Imagine two people trying to order a pizza. Even after the first time they agree on "Pepperoni," the second time they say, "I want the pizza with the red circles and spicy cheese on the dough." They never learned to just say "Pepperoni."
- The Result: They actually got worse at the game over time because they were so busy talking that they lost track of which basket was which. They failed to build that "secret handshake."
3. The Mixed Teams (The Frustrating Partners) 😫
When humans played with AI, it was a disaster for the humans.
- Human as Director (Talking to AI): The human tried to be concise, but the AI kept asking for unnecessary confirmations or repeating things. The human had to do all the heavy lifting.
- AI as Director (Talking to Human): The AI would describe the baskets in a long, confusing way. Sometimes, the AI would describe the wrong basket entirely (a "hallucination"), and the human would try to fix it, but the AI would just ignore the correction and keep talking.
- The Analogy: It's like trying to drive a car with a passenger who keeps changing the destination without telling you. You know where you're going, but the passenger keeps insisting you turn left, even though you're already there.
Why Did the AI Fail?
The paper suggests that AI is like a photocopier of text, not a collaborator.
- Humans learn by interacting. They listen to what their partner needs and adjust.
- AI (in this study) just looked at its instructions and the history of the chat and said, "Okay, I will describe the object again using the most detailed words I know." It didn't realize, "Hey, my partner already knows what this is!"
The AI failed to understand the concept of "Common Ground." It didn't realize that once something is established, you don't need to explain it again. It violated the basic rule of conversation: Don't say too much if you don't have to.
The Big Takeaway
If we want AI to be a good partner for humans (like a doctor's assistant or a co-pilot), it needs to learn how to listen and adapt, not just talk. Currently, even the smartest AI models are terrible at building that shared understanding. They are like a student who memorized the textbook but has never actually had a conversation with a friend.
In short: Humans get smarter together by building a shared language. AI just keeps repeating the dictionary.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.