Implicit vs. Explicit Prompting Strategies for LVLMs in Referential Communication
This paper resolves contradictory findings on LVLMs' ability to coordinate efficient referring expressions by demonstrating that while models succeed under explicit prompting, they fail to infer the need for communicative efficiency from implicit prompts, revealing a fundamental divergence between human and AI communication.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
The Big Question: Do AI Partners Learn to Talk Less?
Imagine you and a friend are playing a game where you have to describe a specific basket from a pile of 18 different baskets so your friend can pick the right one out of their own pile.
In the first round, you might say, "The one with the red handle and a flower pattern on the side."
By the fifth round, if you are human, you and your friend will likely have developed a secret shorthand. You might just say, "The red flower one," or even just "Red flower." You do this because you know your friend understands what you mean. This is called building a "conceptual pact."
Two recent studies looked at whether AI models (specifically Large Vision-Language Models, or LVLMs) can do this same thing.
- Study A said: "No, the AI keeps giving long, boring descriptions every single time, even after five rounds."
- Study B said: "Yes, the AI gets shorter and smarter over time, just like humans!"
This paper by Peter Zeng and colleagues asks: Why did these two studies get opposite answers?
The Detective Work: It's All About the Instructions
The authors set up a controlled experiment to solve this mystery. They realized the two studies weren't just testing different AI models; they were giving the AI different instructions (prompts).
Think of it like giving directions to a taxi driver:
- The "Implicit" Prompt (Study A style): The driver is told, "Drive efficiently and be helpful."
- Result: The driver drives efficiently in terms of safety and rules, but they still take the long scenic route because they weren't told to cut corners. They don't realize that "efficient" means "shortest path" in this specific game.
- The "Explicit" Prompt (Study B style): The driver is told, "Drive efficiently. In fact, I want you to use the absolute shortest route possible. After the first few turns, just give me 1 or 2 words of direction."
- Result: The driver immediately starts taking shortcuts and using slang because they were explicitly told to do so.
What the Experiment Found
The authors ran the basket game with AI partners using both types of instructions.
1. When given the "Implicit" instructions (Just "be concise"):
The AI models were accurate, but they were wordy. They kept describing the baskets in full detail every single time, just like Study A found. They did not figure out on their own that they should start using short codes like "Red flower." They treated every round as a fresh conversation, ignoring the history they had built with their partner.
2. When given the "Explicit" instructions (Just "be concise" + "use 1-2 words later"):
The AI models changed their behavior completely. They started shortening their descriptions, just like Study B found. By Round 5, they were using very short phrases. They looked like they had formed a "conceptual pact."
The Catch:
The authors point out a crucial difference between the AI and humans.
- Humans shorten their speech because they understand their partner and feel a shared connection.
- AI only shortens its speech because it was told to. If you remove the specific command to "use 1-2 words," the AI reverts to being long-winded. It doesn't naturally infer that "we are partners, so we can be brief."
The Conclusion: It's a Script, Not a Spark
The paper concludes that the conflicting results from the previous studies weren't because one AI was "smarter" than the other. It was simply because one study gave the AI a script that forced it to be short, while the other study left it to figure it out on its own.
The Takeaway:
AI can mimic human-like efficiency, but only if you hold its hand and tell it exactly how to do it. It doesn't naturally develop the social intuition to say, "Hey, we've done this before, let's keep it short." It's following orders, not building a relationship.
Summary Analogy
Imagine two students taking a test.
- Student A is told: "Write a good essay." They write a 5-page paper.
- Student B is told: "Write a good essay, but try to keep it under 100 words." They write a 90-word summary.
If you asked, "Can students write short essays?", the answer depends entirely on which instruction you gave them. This paper proves that the AI isn't naturally choosing to be short; it's only short because the teacher (the prompt) told it to be.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.