← Latest papers
🤖 machine learning

Tacit Coordination of Large Language Models

This paper presents the first large-scale evaluation of tacit coordination in Large Language Models, revealing that while LLMs often match or outperform humans in coordinating without communication via focal points, they consistently fail in tasks requiring numerical common sense or cultural nuance, thereby highlighting significant social limitations in their latent notions of salience.

Original authors: Ido Aharon, Emanuele La Malfa, Michael Wooldridge, Sarit Kraus

Published 2026-06-17
📖 5 min read🧠 Deep dive

Original authors: Ido Aharon, Emanuele La Malfa, Michael Wooldridge, Sarit Kraus

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

The Core Problem: The "Blind Date" Dilemma

Imagine you and a friend are blindfolded in a huge, empty stadium. You both have to pick one specific seat to sit in. If you pick the same seat, you win a prize. If you pick different seats, you get nothing. You cannot talk to each other. You cannot signal. You just have to guess where the other person will sit.

In game theory, this is called tacit coordination. Humans are surprisingly good at this. We don’t pick random seats; we pick "salient" ones—seats that stand out. Maybe it’s the very first seat in the front row, or the exact center of the stadium. These standout choices are called Focal Points. They are like social magnets that pull people together without any words being spoken.

The researchers wanted to know: Can AI (Large Language Models, or LLMs) do this? Can an AI guess where a human will sit, or where another AI will sit, without talking?

The Experiment: Testing AI’s "Social Radar"

The team tested over 20 different AI models (including famous ones like Llama, Qwen, and GPT) against human data from previous experiments. They used three main scenarios:

  1. The "Pick a Number" Game (Amsterdam & Nottingham):
    Humans were asked to pick answers from a list. The AI had to guess what the humans picked.

    • The Result: The AI was actually better than humans at some tasks, especially when the answer was obvious (like picking "1" from a list of 1 to 100).
    • The Failure: The AI struggled when the "right" answer depended on culture or subtle context. For example, if the focal point was something culturally specific to Nottingham, England, the AI often missed it. It relied too much on structure (like "pick the middle option") rather than cultural meaning.
  2. The Bargaining Table (Cooperation vs. Selfishness):
    Imagine two players sharing a table with coins on it. They have to agree on who gets which coin. If they disagree, they both lose money.

    • The Result: The AI acted like a very polite, cooperative partner. It tried hard to avoid conflict and maximize the total reward for both sides. It performed similarly to a human who is trying to be fair and cooperative.
  3. Search and Rescue (The Wilderness Test):
    This was the most realistic test. The AI was given a map of a wilderness area where a hiker went missing. It had to predict the single most likely spot to find the hiker.

    • The Result: The AI did well only when the hiker was likely to be at a "focal" location—like a trail junction, a river crossing, or a landmark. If the hiker was just lost in a random patch of woods, the AI’s guesses were no better than chance.
    • Key Insight: The AI isn’t "thinking" about the hiker’s psychology; it’s looking for structural landmarks that stand out on the map, just like a human would.

The Big Surprise: Thinking Harder Doesn’t Help

You might think that if you ask an AI to "think step-by-step" or "reason deeply" about where to coordinate, it would do better.

It doesn’t.

The paper found that forcing the AI to reason more often made it worse at coordination. Instead of finding a clever social cue, the AI would overthink and end up picking an arbitrary option (like "I’ll just pick the first one because it’s first"). Tacit coordination isn’t a logic puzzle; it’s a social instinct. Over-analyzing it breaks the instinct.

The Fix: The "Culture" Cheat Code

The researchers discovered a simple trick to make the AI coordinate better with humans. They didn’t use complex training or new data. They just changed the prompt (the instruction given to the AI).

  • Bad Prompt: "Pick the most logical answer."
  • Good Prompt: "Pick the answer that is most culturally relevant to humans."

When they asked the AI to consider culture, its coordination scores jumped up, often matching or beating human performance. This suggests that AI doesn’t naturally share our cultural "common sense," but it can mimic it if explicitly told to look for cultural cues.

The Takeaway: AI is a Structuralist, Not a Culturalist

The paper concludes with a warning and an insight:

  1. AI is Great at Structure: If the coordination problem has a clear, logical, or structural focal point (like the center of a grid or a major road intersection), AI is excellent at finding it.
  2. AI is Bad at Nuance: If the focal point relies on shared cultural knowledge, inside jokes, or subtle social norms, AI struggles. It defaults to "centrality" (picking the middle) even when humans would pick something else for cultural reasons.
  3. Don’t Assume Shared Mindset: We cannot assume that AI "thinks" like us. It doesn’t have the same cultural background. If we deploy AI in teams with humans (like in search and rescue or business planning), we need to explicitly prompt it to align with human cultural expectations, or it might "coordinate" on the wrong thing.

In short: AI is like a brilliant architect who knows exactly where the center of the building is, but it doesn’t know that humans prefer to meet at the coffee shop on the left because it’s where everyone hangs out. You have to tell the architect about the coffee shop, or it will keep sending everyone to the center lobby.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →