← Latest papers
💬 NLP

Aligned but Not Partner-Specific: Distinguishing How Multimodal LLM Agents Succeed in Reference Games Without Human-Like Conventions

This paper demonstrates that while multimodal LLM agents achieve coordination in reference games through high label alignment, they fail to develop human-like partner-specific conventions, instead relying on consistently verbose descriptions that do not compress over interaction history.

Original authors: Po-Ya Angela Wang, Chinmaya Mishra, Aslı Özyürek, Paula Rubio-Fernández, Esam Ghaleb

Published 2026-06-09
📖 4 min read☕ Coffee break read

Original authors: Po-Ya Angela Wang, Chinmaya Mishra, Aslı Özyürek, Paula Rubio-Fernández, Esam Ghaleb

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

The Big Question: Are AI Partners or Just Echoes?

Imagine you and a friend are playing a game where you have to describe abstract shapes (like a "dinosaur" made of folded paper) to each other without naming them directly.

How Humans Play:
At first, you might say, "That one that looks like an amber-orange colored dinosaur with a long tail." But after you've played a few rounds with the same person, you develop a secret shorthand. You just say, "The Dinosaur." You both know exactly what that means because you built that shared history together. You get faster and use fewer words. This is called forming a "conceptual pact."

How the AI Played:
The researchers asked: Do AI agents (like advanced chatbots) do the same thing? Do they learn to use shorter, secret words with their specific partner, or do they just keep saying the same long, detailed things every time?

The Experiment: The "Fake Partner" Test

To find out, the researchers set up a clever test using a "Pseudo-Dyad" (a fake partner) baseline. Think of it like this:

  1. Real Team: Two humans (or two AIs) play 45 rounds together. They build a history.
  2. The "Fake" Team: The researchers took Round 1 from Team A and paired it with Round 1 from Team B. Then they took Round 2 from Team A and paired it with Round 2 from Team C.
    • The Catch: These "fake" partners have never met. They have no shared history. They are strangers.

If the AI is truly learning a "secret language" with its partner, it should fail when paired with a stranger. If it just uses a standard "dictionary" it learned during training, it will sound the same whether it's talking to a friend or a stranger.

The Results: The "Verbose Robot" vs. The "Efficient Human"

The study found a massive difference between how humans and AI handled the game.

1. The Effort Level (The "Backpack" Analogy)

  • Humans: Imagine carrying a heavy backpack full of words. In the beginning, the backpack is huge. As you play with the same person, you start throwing things out of the backpack because you don't need them anymore. By the end, you are running light and fast. Humans got much more efficient over time.
  • AI: The AI kept the backpack exactly the same size the whole time. It never threw anything away. It kept describing the shapes in incredibly long, detailed sentences (e.g., "A bright amber-orange folded ribbon shape") from the very first round to the very last. It didn't get "lighter" or faster, even though it was playing with the same partner.

2. The "Secret Code" (The "Inside Joke" Analogy)

  • Humans: When humans played with their real partner, they used short, shared words (inside jokes) a lot. But when they played with the fake partner (the stranger), they stopped using those short words and went back to long descriptions. This proved their short words were a result of their specific relationship.
  • AI: The AI used short, shared words almost 100% of the time, regardless of who they were talking to. Whether they were talking to their real partner or a stranger they had never met, they used the exact same vocabulary.

The Conclusion: Coordination Without Connection

The paper concludes that AI agents are aligned but not partnered.

  • Humans succeed by building a unique, compact relationship. They learn to trust each other and drop the extra words.
  • AI succeeds by being exhaustively descriptive. It wins the game not by forming a bond, but by saying so many words that the partner has to understand. It's like a robot that never learns to say "Pass the salt" and instead says, "Please hand me the white cylindrical container containing sodium chloride that is currently on the table to my left." It works, but it's inefficient and doesn't change based on who is listening.

In short: The AI didn't learn to speak its partner's language; it just spoke its own very loud, very detailed language so clearly that everyone understood it, even strangers. It achieved coordination, but it didn't form a convention.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →