Theory of Mind and Persuasion Beyond Conversation: Assessing the Capacity of LLMs to Induce Belief States via Planning and Action
This paper introduces the Non-Conversational Planning Theory of Mind (NCP-ToM) framework to evaluate how well large language models can induce specific belief states through actions rather than conversation, finding that while GPT-5 outperforms humans in success rate on these tasks, it remains less robust across contexts and, like humans, struggles more with inducing false beliefs than true ones.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
The Big Idea: The "Silent Puppet Master" Test
Imagine you are playing a game of chess, but instead of moving pieces to checkmate the king, your goal is to make your opponent believe something that isn't true, or to make them know something that is true, without ever saying a single word. You can't talk to them; you can only move objects around the board or hide things in boxes.
This paper introduces a new way to test AI (Large Language Models or LLMs) called NCP-ToM (Non-Conversational Planning Theory of Mind).
- Theory of Mind (ToM): This is the human ability to understand what other people are thinking or believing.
- Non-Conversational: The AI isn't allowed to chat, persuade, or explain. It has to use actions to change what others believe.
The researchers wanted to see if AI agents could act like a "silent puppet master," arranging the world so that other characters end up with specific beliefs, just by moving things around.
The Setup: A Digital Game of "Hide and Seek"
To test this, the researchers built a digital playground based on a classic psychology test called the Sally-Anne test.
- The Classic Test: In the original version, you watch a video where Sally puts a marble in a basket and leaves. Anne moves the marble to a box. You are asked, "Where does Sally think the marble is?" (The answer is the basket, because she didn't see the move). This is a Question & Answer test. You just watch and guess.
- The New Test (NCP-ExploreToM): In this paper, the AI isn't just watching; it is inside the game. The AI is given a goal, like: "Make Sally believe the marble is in the basket, even though it's actually in the box."
- The AI's Job: The AI has to figure out the steps to make this happen. It might have to:
- Move Sally out of the room.
- Move the marble to the box.
- Leave Sally out of the room so she doesn't see the switch.
- Bring Sally back in.
If the AI does this correctly, Sally will have a "false belief" (she thinks it's in the basket). If the AI fails, Sally might see the switch and know the truth.
Who Played the Game?
The researchers tested:
- Six Top-Tier AI Models: Including GPT-5, Gemini 2.5 Pro, and various versions of Claude.
- 40 Human Participants: Real people playing the same game on a computer.
They ran 600 different scenarios (tasks) with different levels of difficulty, different rooms (like a hospital, a hotel, or a wedding), and different types of goals (making someone believe something true vs. making them believe something false).
The Results: The Scoreboard
Here is what happened when they compared the players:
1. The "False Belief" Challenge is Hard
Just like humans, the AI models found it much harder to trick someone into having a false belief (making them think something that isn't true) than to help them have a true belief (making them know the truth).
- Analogy: It's easier to show someone a picture of a cat so they know there's a cat (True Belief) than to hide the cat and make them think it's a dog (False Belief).
- Why this matters: The paper notes this is a "positive signal." It suggests AI might be better at helpful tasks (teaching, assisting) than deceptive ones (tricking people).
2. The "GPT-5" Surprise
The newest model, GPT-5, was the superstar.
- It solved about 80% of the tasks.
- It was the only model that actually did better than the human participants.
- However, humans were more consistent. The AI's performance dropped significantly depending on the "setting" (e.g., it did great in a government building scenario but worse in a wedding scenario). Humans were steady no matter the setting.
3. The "Planning" Gap
The researchers compared two ways of testing the AI:
- The "Quiz" (Q&A): "Here is a story. What does Sally think?"
- The "Action" (Agentic): "Here is a goal. Go do it."
Usually, people assume the "Action" test is much harder than the "Quiz." For most AI models, this was true. But for GPT-5, the gap disappeared. It performed almost as well on the "Action" test as it did on the "Quiz." This suggests that for the smartest models, actually doing the planning is becoming just as easy as just talking about it.
4. Complexity Kills Performance
As the tasks got more complicated (requiring the AI to track multiple people's beliefs at once, or making three different people believe three different things), everyone's score went down. This is like trying to juggle more balls; eventually, you drop one.
What Does This Mean? (According to the Paper)
The paper concludes with a few key takeaways, sticking strictly to what they found:
- AI is getting good at "Silent Persuasion": Modern AI can plan a sequence of physical actions to change what other characters believe, without saying a word.
- Safety Check: The fact that AI struggles more with "false beliefs" (deception) than "true beliefs" is good news for safety. It suggests they aren't naturally inclined to lie or manipulate via action, though they can do it if asked.
- The "Context" Problem: AI is still a bit fragile. If you change the story from a "hospital" to a "wedding," the AI might get confused. Humans don't get confused by these changes.
- The "Quiz" Trap: We can't just assume that if an AI is good at answering questions about a story, it will be good at acting out that story. For the very best models (like GPT-5), the old rule that "Quizzes are easier than Actions" might no longer be true.
The Bottom Line
This paper shows that AI is evolving from a "chatbot" that answers questions into an "agent" that can plan and act to shape reality. While the smartest AI can currently outperform humans at these specific logic puzzles, it still lacks the human ability to stay consistent across different real-world situations. The researchers warn that while this is a cool capability for helpful assistants, it also means we need to be careful about how we evaluate AI safety, because the old tests (just asking questions) might not be enough anymore.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.