← Latest papers
🤖 AI

MINDGAMES: A Live Arena for Evaluating Social and Strategic Reasoning in Multi-Agent LLMs

This paper introduces Mindgames, a multi-game evaluation arena and dataset comprising 29,571 interactions from a 2025 competition, designed to assess large language models' social and strategic reasoning capabilities while revealing critical limitations such as brittle rule adherence and environment-specific leaderboard validity.

Original authors: Kevin Wang, Anna Thöni, Benjamin Kempinski, Bobby Cheng, Jianzhu Yao, Benjamin Finch, Leon Guertler, Viraj Nadkarni, Yihan Jiang, Aliaksei Korshuk, Alexander Buyantuev, Ilya Makarov, Siyuan Wu, Yu-Chi
Published 2026-05-29
📖 6 min read🧠 Deep dive

Original authors: Kevin Wang, Anna Thöni, Benjamin Kempinski, Bobby Cheng, Jianzhu Yao, Benjamin Finch, Leon Guertler, Viraj Nadkarni, Yihan Jiang, Aliaksei Korshuk, Alexander Buyantuev, Ilya Makarov, Siyuan Wu, Yu-Chi Cheng, Yan-Ru Ju, Ti-Rong Wu, I-Hsuan Chu, Yu-Yu Yang, I-Chen Wu, Yitian Huang, Qinlu Cao, Yiheng Sun, Yuhong Dai, Hongkun Yao, Jingxuan Fu, Jiwei Zhang, Hao Liao, Mossimo Ebeling, Govind Arun, Sadhvik Bathini, Mihir S Arya, Avinash Anish, Aditya Ranjan, Kirtana Sunil Phatnani, Paval KS, Vrushali Mehta, Aravind S, Nikhil Arora, Tanya Upadhyay, Amol Bandagale, Yuan Lu, ChunEn Hsiao, YuTing Lin, Arvin Chung, Jerry John Thomas, Mathieu Laurière, Leshem Choshen, Yoram Bachrach, Pramod Viswanath, Maria Polukarov, Cheston Tan, Tal Kachman, Atlas Wang

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine a giant, digital playground where artificial intelligence (AI) agents don't just play solitaire alone, but are forced to interact, lie, negotiate, and strategize against each other in real-time. This is MINDGAMES, a new competition and research project designed to test how well Large Language Models (LLMs) can understand other people's minds—a skill psychologists call "Theory of Mind."

Here is the story of the paper, broken down into simple concepts.

1. The Problem: AI is Good at Tests, Bad at People

Until now, we tested AI by giving them static riddles or single-shot questions (like "If Alice thinks Bob is lying, what does she do?"). The paper argues this is like testing a swimmer by having them stand in a pool and kick their legs. It doesn't tell you if they can actually swim in a rough ocean.

Real life is messy. It involves:

  • Hidden information: Not knowing what the other person knows.
  • Deception: Lying to get ahead.
  • Cooperation: Working together when you have different goals.
  • Adaptation: Changing your strategy because your opponent changed theirs.

The authors built MINDGAMES to be that "rough ocean." It's a live arena where 944 AI agents from 76 different teams competed in four specific games.

2. The Arena: Four Games, Four Different Skills

Think of these games as four different gyms, each training a different muscle of social intelligence:

  • Colonel Blotto (The Chess Match): Two players split a fixed number of troops across three battlefields. You don't know where the opponent is putting their troops.
    • Skill tested: Opponent Modeling. Can you guess what the other person is thinking and outsmart them without talking?
  • Iterated Prisoner's Dilemma (The Trust Exercise): Three players chat freely, then decide to cooperate or betray each other for points.
    • Skill tested: Trust and Betrayal. Can you build a reputation, detect a liar, and stick to a promise?
  • Codenames (The Secret Handshake): Two teams of two. One person (the Spymaster) gives a one-word clue to help their partner guess secret words on a board, without accidentally revealing the "Assassin" word.
    • Skill tested: Shared Understanding. Can you communicate complex ideas through tiny, constrained hints?
  • Secret Mafia (The Detective Game): Six players. Some are "Mafia" (knowing each other secretly), and some are "Villagers" (trying to find the Mafia). They debate and vote to eliminate suspects.
    • Skill tested: Deception and Deduction. Can you lie convincingly while figuring out who is lying?

3. The Results: The "Error-Survival" Trap

The competition was a huge success in gathering data (nearly 30,000 games!), but it revealed a surprising flaw in how we usually rank AI.

The "Clean" Games:
In games like Colonel Blotto and Prisoner's Dilemma, the rankings made sense. The AI that played the best strategy won. It was like a clear race.

The "Messy" Games:
In Secret Mafia and Codenames, things got weird. The paper found a major problem called the "Error-Survival Confound."

Imagine a game of musical chairs where the music stops, and everyone has to sit. If you are a great dancer but you trip and fall, you lose. But in these AI games, many players were so bad at following the rules that they tripped constantly.

  • In Secret Mafia, the "winners" weren't necessarily the best liars or detectives. They were just the ones who didn't trip over their own feet as often as everyone else.
  • Because the game is so complex, many AIs made "fatal errors" (like revealing their secret role or voting illegally) and got kicked out early. The agents that survived simply by not crashing were ranked higher than the agents that were actually smarter but made one small mistake.

The Lesson: In complex social games, a leaderboard might just be measuring "who is the least clumsy" rather than "who is the smartest."

4. How the Winners Did It

The paper analyzed the top-performing teams and found they didn't just rely on having a "bigger brain" (more computing power). Instead, they used clever engineering tricks:

  • The "Scaffolding" Approach: Instead of just asking the AI "What do you do?", the winners gave the AI a checklist, a memory notebook, and a code interpreter to do the math. It's like giving a student a calculator and a study guide instead of just a textbook.
  • Training vs. Prompting: For smaller AI models, training them on past game data worked best. For the biggest, most powerful models, prompting them with clever instructions (without retraining) worked better.
  • The "Don't Lie" Problem: The AI agents struggled to lie. Because they are trained to be helpful and truthful, they often accidentally revealed they were "Mafia" when they tried to bluff. The best agents had to be explicitly taught how to break this "truthful" habit.

5. The Toolkit for the Future

The authors didn't just publish the results; they handed everyone the keys to the playground. They released:

  • The Dataset: A massive library of 29,000 games with every move and thought logged, so other researchers can study how AI thinks (or fails).
  • MG-Ref (The Reference Set): A "frozen" pool of the best, most stable players from the competition. New AI agents can now play against this specific group offline to get a fair score without needing a live server.
  • A New Way to Measure: They suggest that future tests shouldn't just look at the final score. They must also report how many errors the AI made. If an AI wins because everyone else crashed, that's not a real victory.

Summary

MINDGAMES is a reality check for AI. It shows that while AI is getting better at playing games, it is still very fragile. It struggles with the messy, long-term, deceptive nature of human social interaction. The paper concludes that to truly measure "social intelligence," we need to stop just looking at who wins the game and start looking at how they played, how many mistakes they made, and whether they survived because they were smart or just because everyone else was clumsy.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →