← Latest papers
🤖 AI

Improvisational Games as a Benchmark for Social Intelligence of AI Agents: The Case of Connections

This paper introduces "Connections," an improvisational wordplay game designed to benchmark the social intelligence of AI agents by evaluating their ability to integrate knowledge retrieval, summarization, and the assessment of other agents' cognitive states within a constrained collaborative environment.

Original authors: Gaurav Rajesh Parikh, Angikar Ghosal

Published 2026-04-02
📖 4 min read☕ Coffee break read

Original authors: Gaurav Rajesh Parikh, Angikar Ghosal

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are playing a game of charades, but with a twist: you can't act, you can only speak in riddles, and your teammates are trying to guess a secret word while a "referee" tries to stop them.

This paper introduces a new way to test how "socially smart" AI is. Instead of asking an AI to solve a math problem or write a poem, the researchers put it in a game called Connections.

Here is the breakdown of the game and what the researchers learned, explained simply.

The Game: "Connections"

Think of this as a high-stakes game of "20 Questions" mixed with a word puzzle.

  • The Setup: One player (the Setter) picks a secret word (like "Catamaran"). The other players (the Guessers) don't know the word.
  • The Clue: The Setter reveals the first letter (e.g., "C").
  • The Dance: A Guesser says a clue, like "A rug that sometimes flies."
    • If the other Guessers say "Carpet!" at the same time, they get a point, and the next letter is revealed ("A" -> "CA").
    • The Catch: The Setter is listening. If the Setter figures out the word before the Guessers do, they can shout "Blocked!" and stop the Guessers from getting the letter.
  • The Goal: The Guessers win if they guess the whole word before running out of clues. The Setter wins if they can block them.

Why is this hard for AI?

Most AI tests check if a robot can remember facts or do math. This game checks if an AI can read the room.

To win, an AI Guesser has to do three tricky things at once:

  1. Be Clever: Come up with a clue that is smart enough to make you think of the word, but not so obvious that the Setter guesses it immediately.
  2. Be a Mind Reader: You have to guess what your teammates know. If you give a clue about "Medical Terms," but your teammate is a farmer, they won't get it. You need to find a clue that both of you know.
  3. Be a Spy: You have to guess what the Setter doesn't know. If the Setter is a physics genius, don't use a physics clue; they will block it instantly.

The Experiment: Robots Playing Robots

The researchers set up a game with three AI agents (using a powerful model called GPT-4o).

  • Agent 0 was the Setter.
  • Agent 1 & 2 were the Guessers.

They played hundreds of rounds. Here is what they found:

1. The "Same Brain" Problem
At first, the game was boring. Because all the AIs were built from the same "brain" (the same training data), they all thought the same way.

  • Analogy: Imagine three twins playing charades. If one twin says "A red fruit," the other twins instantly know it's an apple, but the Setter twin also knows it's an apple and blocks it immediately. The AIs kept getting blocked because they all had the exact same "mental dictionary."

2. The "Social" Breakthrough
The researchers tried something new: they gave the AIs different "personalities" (like "You are a 60-year-old carpenter" vs. "You are a 20-year-old gamer").

  • The Result: The AIs started to get better. They began to realize, "Oh, my teammate knows about video games, but the Setter doesn't. I should use a gaming clue!"
  • This showed that the AI could start to simulate what another person knows, which is a huge step toward true social intelligence.

The Big Takeaway

The paper argues that for AI to be truly intelligent, it can't just be a super-smart encyclopedia. It needs to be a social chameleon.

  • Current AI: "I know the definition of 'Catamaran'."
  • Socially Smart AI: "I know that my friend knows 'Catamaran' because we went to the beach together, but the referee doesn't, so I'll give a clue about boats that my friend will get but the referee won't."

Why Does This Matter?

The authors believe that games like this are the future of testing AI. Real life isn't just about answering questions correctly; it's about understanding what others know, what they don't know, and how to communicate effectively with them.

If an AI can learn to play this game well, it means it's taking its first steps toward understanding human culture, context, and the messy, unpredictable nature of human conversation. It's moving from being a calculator to being a teammate.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →