Auditing Belief-Conditioned LLM Agents in Hidden-Information Social Deduction Games
This paper introduces an auditable framework for evaluating LLM agents in hidden-information social deduction games, demonstrating that while active belief maintenance significantly improves good-side win rates in Werewolf, the underlying mechanism remains unresolved due to low direct action-belief consistency and unexpected benefits when beliefs are restricted to werewolves.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine a high-stakes game of Werewolf, but instead of a group of friends sitting around a table, you have nine super-smart AI bots playing against each other. In this game, some bots are "Good Guys" (villagers, a seer, a witch) and some are "Bad Guys" (werewolves). The catch? The werewolves know who their teammates are, but the good guys have to guess who the monsters are just by listening to what everyone says. The problem is, the werewolves are liars, and the good guys are often confused.
For a long time, researchers tried to see if these AI bots could get better at the game just by playing more rounds and looking at who won. But that's like trying to learn how to drive by only looking at the final destination. If you crash, you don't know if it was because you forgot to brake, the road was icy, or your passenger yelled at you. The final result is too messy to tell you why the AI made a bad move.
The Detective's Notebook: A New Way to Watch
The authors of this paper built a special "detective's notebook" for these AI agents. They created a system where an external observer (the notebook) keeps a running list of who it thinks is a werewolf based on the evidence, even though the AI bots themselves don't see this list.
Think of it like a shadow coach standing behind the AI. The coach whispers, "Hey, Player 5 looks suspicious," but the AI is free to ignore that whisper and do whatever it wants. The system records every time the AI listens to the coach and every time it ignores the coach. This turns the AI's messy, hidden thoughts into a replayable video game where we can pause, rewind, and ask, "Wait, why did you vote for Player 3 when the evidence pointed to Player 5?"
The Big Surprise: The Coach Helps, But Not How You Think
The researchers ran 1,080 of these games to see what happened. They compared two groups:
- The "No-Coach" Group: The AI played without any extra hints.
- The "Active-Coach" Group: The AI got to hear the shadow coach's suspicions.
Here is the cool part: The "Active-Coach" group won much more often. When they played 200 games against the "No-Coach" group using the exact same starting conditions, the "Good Guys" win rate jumped from 0.205 (about 20%) to 0.390 (almost 40%). That's a huge leap!
But here is the twist, and it's the most important part of the story: The paper explicitly says we don't know why this happened.
You might think, "Oh, the AI just listened to the coach and got smarter!" But the data says no.
- The AI only followed the coach's top recommendation about 0.21 (or 21%) of the time. That's barely one in five!
- In fact, when they gave the "Coach" only to the werewolves (the bad guys), the Good Guys actually won more often than when they gave the coach only to the Good Guys. This is backwards! If the coach was just helping the person holding it, the Good Guys should have won more when they had the coach. Since they didn't, the paper argues that the "Coach" isn't just a simple tool for better guessing.
So, the paper rules out the idea that the AI simply got better at "reading the room" because it had the notes. The mechanism is a mystery. The authors suggest the "Coach" might be changing the AI's behavior in weird, indirect ways, or maybe the presence of the notes changes how the AI thinks, even if it doesn't read them.
Catching the Mistakes Before They Happen
Even though we don't know the full "why," the detective notebook did something amazing: it caught a specific, dangerous mistake.
In Werewolf, the "Witch" has a potion that can kill someone. If she poisons a Good Guy by mistake, it's a disaster.
- Without the Coach: The Witch tried to poison someone in 29 games, and she was wrong 22 times (a 76% failure rate).
- With the Coach: The Witch tried to poison someone in only 11 games, and she was wrong only 4 times (a 36% failure rate).
The notebook helped the Witch stop being so reckless. It didn't make her perfect, but it stopped her from making the worst possible mistakes.
What the Paper Says We Should NOT Do
The researchers also tested a very logical idea: "If the AI isn't listening to the coach enough, let's just force it to listen!" They tried to make the AI obey the coach's notes more strictly.
It failed.
When they forced the AI to follow the coach, the win rate didn't go up. In fact, it went down slightly (from 0.412 to 0.362). The paper explains that sometimes the coach's notes are just "meh"—the evidence is weak, and the coach isn't sure who the werewolf is. Forcing the AI to follow a weak guess is like forcing a driver to turn left when the road sign is blurry; it just leads to more crashes. The paper concludes that you can't just force the AI to obey; it needs to know when the coach is actually confident.
The Bottom Line
This paper doesn't claim to have "solved" Werewolf or made the AI a genius. Instead, it built a super-powerful microscope.
- What it proved: Giving the AI a "shadow coach" is associated with better results and fewer terrible mistakes.
- What it ruled out: It proved that simply forcing the AI to follow the coach doesn't work. It also ruled out the idea that the win rate jump was just because the computer was running faster or slower (they checked the "load" and found no difference).
- What is still a mystery: Why does the win rate jump so high if the AI barely listens to the coach? The paper says the "why" is still unresolved.
The main takeaway is that in games where secrets are hidden and lies are common, you can't just look at who won. You need a way to watch how the AI thinks, record its doubts, and replay its mistakes. The "detective notebook" turns a black box into a replayable video, making it safer to experiment and learn, even if we don't fully understand the magic behind the curtain yet.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.