Deceive, Detect, and Disclose: Large Language Models Play Mini-Mafia
This paper introduces "Mini-Mafia," a simplified social deduction game and benchmark that demonstrates how a compact analytical formula based on intrinsic deception, detection, and disclosure parameters can accurately predict large language model interactions and collective outcomes across diverse model combinations.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine a game of Mafia (also known as Werewolf), but stripped down to its absolute bare bones. That's what this paper introduces: Mini-Mafia.
Think of the original game as a complex, chaotic city with hundreds of people, secret meetings, and long nights. The researchers realized that to truly understand how AI agents interact, they needed a simpler "laboratory." So, they built a tiny, four-player version where the rules are fixed, the night phase is automatic, and the whole game boils down to one critical conversation during the day.
Here is the simple breakdown of what they did and what they found:
1. The Setup: A Three-Person Standoff
In this mini-game, there are only four players, but only three survive to talk:
- The Mafioso: The liar. Their job is to convince everyone they are innocent.
- The Detective: The truth-teller. They know who the Mafioso is and must convince the others to vote them out.
- The Villager: The judge. They have no secret information; they just have to listen to the other two and decide who to trust.
The game ends after one round of talking and one vote. If the Villager votes for the Mafioso, the town wins. If they vote for the Detective (or tie), the Mafioso wins.
2. The Big Discovery: A Simple Math Formula
The researchers played this game thousands of times using 10 different Large Language Models (LLMs) as the players. They expected the results to be a messy black box where you just see who won and who lost.
Instead, they found a simple mathematical recipe that predicts the outcome almost perfectly.
Imagine the game's result is determined by a tug-of-war:
- Deception (): How good the Mafioso is at lying.
- Disclosure (): How good the Detective is at telling the truth.
- Detection (): How good the Villager is at spotting the difference.
The formula they found is: Win Rate = (Deception − Disclosure) × Detection.
Think of it like a chemical reaction:
- If the Mafioso is a great liar and the Detective is bad at explaining the truth, the Mafioso wins.
- If the Detective is great and the Mafioso is bad, the Detective wins.
- But the Villager is the multiplier. If the Villager is "asleep" (low detection), it doesn't matter how good the other two are; the result is random. If the Villager is "awake" (high detection), the gap between the liar and the truth-teller decides the winner instantly.
3. The Benchmark: Testing the AI "Personalities"
Using this formula, the researchers created a "report card" for 10 different AI models. They didn't just ask, "Who is the smartest?" Instead, they asked: "Who is the best liar? Who is the best truth-teller? Who is the best judge?"
The results were surprising:
- The Underdogs Won: Smaller, cheaper models often beat the massive, expensive "flagship" models.
- Grok 3 Mini was the best "Judge" (Villager). It was incredibly good at spotting the liar.
- GPT-5 Mini was the best "Truth-Teller" (Detective). It was the most effective at revealing the truth.
- The Giants Struggled: Some of the most famous, powerful models performed poorly.
- Claude Sonnet 4 was the worst "Judge." It was so bad at detecting the liar that it was basically guessing randomly, even though it's a top-tier model.
4. Why This Matters (According to the Paper)
The paper argues that we usually treat AI interactions like a mystery box: we run a simulation and see what happens. This research shows that we can actually measure specific skills (lying, truth-telling, judging) using simple math.
They also found some funny, human-like quirks in the AI:
- Name Bias: The AI trusted the name "Bob" slightly more than "Diana," making it harder to vote Bob out if he was the liar.
- Recency Effect: If the Detective spoke last before the vote, they had a huge advantage. If they spoke first, people forgot what they said.
Summary
The paper introduces a tiny, simplified game called Mini-Mafia to study how AI agents interact. They discovered that the outcome of these interactions isn't random chaos; it follows a clean, predictable math formula based on three skills: Deception, Disclosure, and Detection.
By measuring these skills, they found that smaller AI models can sometimes outperform giant ones in social situations, and that even AI can be influenced by simple things like the order in which people speak or the names they use. This gives researchers a new, clear way to test and understand AI behavior without needing to run endless, complicated simulations.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.