Triadic Werewolf: A Jester Role for Multi-Hop Theory of Mind in LLMs
This paper introduces the "Triadic Werewolf" game, a multi-hop Theory-of-Mind benchmark featuring a Jester faction with inverted incentives that reveals how current large language models struggle with complex three-way strategic reasoning, often failing to distinguish between genuine suspicion and intentional deception despite self-learning mechanisms.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
The Big Idea: Breaking the "Two-Team" Habit
Imagine you are playing a game of Werewolf (also known as Mafia). Usually, there are two teams: the Villagers (good guys) and the Werewolves (bad guys). The Villagers try to guess who the Werewolves are, and the Werewolves try to hide.
In this setup, if someone acts weird or suspicious, the Villagers usually vote them out. It's a simple rule: Suspicious = Bad.
The researchers in this paper asked a question: Are AI models (LLMs) actually thinking about what other players are thinking, or are they just following a simple rule?
To find out, they added a third, tricky character to the game: The Jester.
The New Character: The Jester
The Jester is a role with a very strange goal: The Jester wins ONLY if the Villagers vote them out.
Think of the Jester like a suspicious actor who wants to get fired.
- The Villagers want to fire the Werewolves.
- The Werewolves want to hide and fire the Villagers.
- The Jester wants to look so suspicious that the Villagers fire them.
This creates a "Triadic" (three-way) problem. Now, when a player looks suspicious, the Villagers have to ask a much harder question:
"Is this person a Werewolf (who I should vote out), or is this person a Jester (who I should not vote out, because voting them out makes me lose)?"
The Experiment: Testing the AI's Brain
The researchers ran 60 games using three different powerful AI models (GPT-4.1, DeepSeek-V3.1, and Llama-3.3-70B). They played with and without a "self-learning" feature, where the Jester could read notes from previous games to get smarter.
Here is what they found, explained through metaphors:
1. The "Cheat Code" Problem (Cue-Sufficiency)
In the old two-team games, the AI could win by just looking for "suspicious" behavior. It was like a student who memorized the answer key: If the teacher frowns, the answer is wrong.
- The Result: When the Jester was added, the AI models failed. They couldn't stop following the "suspicious = bad" rule.
- The Jester's Victory: Because the AI kept voting out the "suspicious" Jester, the Jester won 60–70% of the games. The AI was so good at spotting "suspicion" that it accidentally helped the Jester win.
2. The Werewolves' Self-Sabotage
The Werewolves (the actual bad guys) were supposed to protect the Jester from being voted out, because if the Jester leaves, the Werewolves lose too.
- The Result: The AI Werewolves were terrible at this. In 60–70% of the games, the AI Werewolves voted to kick out the Jester on Day 1.
- The Metaphor: It's like a spy team voting to fire their own double-agent because the double-agent was acting "too suspicious," not realizing that firing the double-agent actually helps the enemy win. The AI couldn't connect the dots: Suspicious Person + Jester Goal = Don't Vote Them Out.
3. The "Self-Learning" Loop
The researchers let the Jester read a "cheat sheet" of lessons learned from previous games.
- DeepSeek-V3.1: This model was the smartest. It learned to look suspicious without looking too suspicious. It realized it needed to trick the Villagers while avoiding the Werewolves. It got much better at winning.
- GPT-4.1: This model was already so good at spotting "suspicion" that the cheat sheet didn't help. In fact, it got slightly worse because it was already at the "ceiling" of what it could do with simple rules.
- Llama-3.3: It improved a little, but not as much as DeepSeek.
The Deep Dive: What Was the AI Actually Thinking?
The researchers looked at the "notes" the AI Jesters wrote to themselves to see how they were thinking. They measured "Theory of Mind" (the ability to understand others' minds) in three levels:
- Level 1: "I need to act weird." (Simple)
- Level 2: "I need to act weird so the Villagers think I'm a Werewolf." (Better)
- Level 3: "I need to act like a Werewolf to the Villagers, but not like a Werewolf to the Werewolves, so they don't protect me." (Complex)
The Finding: Only DeepSeek-V3.1 consistently reached Level 3. It understood that it had to play two different roles at once: a "bad guy" for the Villagers and a "normal guy" for the Werewolves. The other models mostly stayed at Level 1 or 2.
The Conclusion
The paper concludes that current AI models are very good at spotting patterns (like "suspicious behavior"), but they struggle with multi-hop reasoning.
- Simple Analogy: Imagine a dog trained to sit when you say "Sit." If you say "Sit" while holding a treat, the dog sits. But if you say "Sit" while holding a different treat that means "Stay," the dog gets confused.
- The Paper's Claim: The AI models are like that dog. They see "Suspicious" and immediately vote "Out." They fail to realize that in this specific game, "Suspicious" can mean "Vote Out" (if it's a Werewolf) OR "Keep" (if it's a Jester).
The researchers argue that to truly test if an AI has a "Theory of Mind," we need games with three competing goals (like this Jester game), not just two. The current "two-team" games are too easy because the AI can win by just following simple rules without truly understanding the other players' hidden motives.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.