MafiaScope: Non-Invasive, Time-Resolved Belief Probing for LLM Agents in Social Deduction Games
MafiaScope is an open testbed that enables non-invasive, time-resolved probing of LLM agents' private beliefs during social deduction games, revealing significant calibration errors and providing tools for visualizing and counterfactually replaying agent reasoning trajectories.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine a group of AI agents playing a high-stakes game of Mafia. In this game, a few players are secret "Mafia" members who know each other and try to kill the innocent "Villagers" at night, while the Villagers try to figure out who the killers are during the day. Usually, when we watch AI play, we only see the final scoreboard: who won, who got voted out, and who lied. It's like watching a magic show and only seeing the final trick, with no idea how the magician's hands moved.
MafiaScope is a new tool that acts like a magical "mind-reading headset" for these AI agents. But here's the catch: it doesn't change the game. It's non-invasive. Every time an AI agent says something out loud to the group, the system quietly pauses, asks the agent a secret question about what it actually believes, and then immediately deletes that answer so it never influences the game again.
Think of it like a reality show where, after every dramatic confession, the camera cuts to a private room to ask the contestant, "Okay, who do you really think is the killer right now?" The contestant answers, the camera records it, and then the answer is locked away in a vault. The contestant never hears their own secret answer, so they can't change their behavior based on it. This lets researchers see the agent's true thoughts versus its public lies.
What Did They Actually Find?
The researchers ran 32 games using a specific AI model (DeepSeek) and collected over 13,800 of these secret answers. Here is what the data revealed:
- The "Spotlight" Delusion: The innocent agents are terrible at guessing what others think of them. They believe they are being suspected 1.5 times more often than they actually are. It's like walking into a party and feeling like everyone is staring at your weird haircut, when in reality, nobody noticed. The paper suggests this is a "machine version" of a human psychological quirk called the "spotlight effect." Interestingly, the actual Mafia agents were much better at this; they knew exactly when they were safe and when they were in trouble.
- Confidence is a Bluff: When the agents said they were "very confident" in their guesses, they were often wrong. The paper measured this using a metric called Expected Calibration Error, which came out to 0.17. In plain English, if an agent says, "I'm 80% sure this person is the Mafia," they are only right about 54.6% of the time. They are overconfident bluffers, not accurate predictors.
- The Truth is a Rollercoaster: The agents' beliefs didn't settle down into a steady truth. Instead, their suspicion levels were volatile. On average, an agent's "top suspect" changed in 48.7% of the secret check-ins. They were circling around the truth rather than locking onto it.
- Lying Works (Sometimes): The Mafia agents won 31 out of 32 games in this specific study. The tool showed that the Mafia successfully kept the Villagers confused, even though the Villagers were getting better at guessing as the game went on (their accuracy rose from 47.6% in round 0 to 75.4% in round 3).
What They Explicitly Ruled Out
The paper is very careful not to overhype these results.
- It's not a "Solved" Problem: The authors explicitly state that these findings describe this specific setup with this specific AI model. They do not claim that all AI agents are overconfident or that all AI Mafia players are this good.
- It's not "Mind Reading" in the Philosophical Sense: The paper argues against the idea that these agents have deep, human-like "beliefs." Instead, they treat the answers as "operational reports"—basically, the AI is just predicting what it would say if asked, not necessarily revealing a soul.
- It's not a Guarantee of Truth: The paper notes that because the agents are self-reporting, they might be lying to the probe itself. However, they found that the agents' secret answers matched their public voting behavior closely (in 64.9% of innocent votes, the agent voted for its top secret suspect), suggesting the reports are at least consistent with their actions.
How Sure Are We?
The paper uses the word "suggests" quite a bit, especially regarding why the agents behave this way.
- The "Pivotal Moment" Experiment: The researchers tried to see if changing just one sentence in the game would change the outcome. They ran a simulation where they "forked" the game (created 30 parallel timelines) to see what would happen if a specific player hadn't said a specific thing.
- They found that fixing one specific lie (by a player named Finley) suggested it might have changed the outcome, raising the chance of a Villager getting eliminated from 0.4 to 0.8.
- However, the paper admits this is not statistically proven. With only 5 reruns per scenario, the difference wasn't big enough to be 100% sure (the statistical p-value was roughly 0.52, which is not significant). They are showing that the tool works to find these moments, not that they have definitively solved the mystery of that specific game.
The Big Picture
MafiaScope isn't a magic wand that fixes AI. It's a microscope. Before this, we only saw the AI's final vote. Now, we can see the messy, confused, overconfident, and sometimes brilliant thought process happening between the votes.
The tool allows researchers to watch the "belief trajectory"—how an agent's mind changes second-by-second. It shows that while AI can play the game, it often doesn't "know" what it knows. It guesses, it gets overconfident, and it sometimes thinks everyone is watching it when they aren't.
The paper concludes by releasing the engine, the visualizer, and the data for anyone to play with. It's an open invitation to watch the AI's mind in real-time, proving that in the world of AI, what you say and what you think are often two very different things.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.