Do we carry the false belief that LLMs have a theory of mind?
Although Large Language Models often achieve high scores on Theory of Mind assessments, this study reveals that they rely on cost-effective heuristics rather than genuinely tracking mental states, indicating they lack robust Theory of Mind capabilities.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Human beings possess a unique cognitive superpower that allows us to navigate the social world: the ability to understand that other people have minds of their own, filled with thoughts, beliefs, and desires that may differ from our own. Psychologists call this capacity "theory of mind." It is the mental machinery that lets us realize that a friend might be looking for something in a place where they last saw it, even if we know that the object has since been moved to a different spot. This skill is not just about being clever; it is the foundation of empathy, communication, and cooperation. For decades, researchers have used a simple story, often involving two characters named Sally and Anne, to test whether children have developed this ability. In the story, Sally hides a toy in a basket and leaves the room. While she is gone, Anne moves the toy to a box. When Sally returns, the question is simple: where will she look for her toy? A child with a developed theory of mind knows Sally will look in the basket, because that is where she believes the toy to be, despite the reality that it is now in the box.
Recently, a new kind of intelligence has emerged in our daily lives: large language models. These are computer programs trained on vast amounts of text that can write, chat, and solve problems with a fluency that often mimics human conversation. Because these models can discuss complex topics and seem to understand context, many researchers and observers have begun to wonder if they, too, possess a theory of mind. If a machine can correctly predict what a fictional character believes, does it truly understand the concept of belief, or is it merely guessing based on patterns it has seen before? This question has become urgent as these tools become more integrated into our social and professional lives, leading to a pressing need to determine if they are genuine social partners or just very sophisticated mimics.
A recent study set out to answer this question by putting five of the most advanced language models to the test using the classic Sally-Anne scenario and several of its variations. The researchers, led by Nicholas Griffen from Université Paris Cité, did not simply ask the models the standard question. Instead, they designed a series of challenges to see if the models could handle the subtle details of the story that require genuine reasoning rather than just memorization. They tested models from major companies, including GPT-5.5 Luna, Claude Sonnet 5, Gemini 3.1 Pro, Grok 4.5, and Mistral Medium 3.5. To ensure the results were fair, the researchers created new, unique versions of the story for each test, changing names, objects, and specific details so the models could not simply recall a previous answer from their training data. They also introduced tricky twists, such as making the containers transparent so the characters could see the toy, or changing the words used to describe where the toy was placed, to see if the models could adjust their reasoning based on new visual or linguistic information.
The results painted a picture of a system that is highly capable in some areas but fundamentally limited in others. When the models were given the standard version of the story, they performed remarkably well, with most answering correctly nearly every time. In fact, when averaged across all the different types of questions, the models scored well above what would be expected by random chance. The top performer, Gemini 3.1 Pro, got the correct answer 84 percent of the time overall. This success suggests that these models have learned the general pattern of the "false belief" task from the vast amount of text they have read. They can recognize the structure of the story and provide the answer that a human would give in a textbook scenario.
However, the study revealed that this success is fragile. When the researchers introduced variations that required the models to use common sense or integrate visual information, the performance dropped dramatically. In one variation, the story described the containers as transparent, meaning the character who returned to the room could actually see the toy in its new location. In a real-world situation, a person would understand that because the character can see the toy, they no longer hold a false belief; they know the truth. The models, however, largely failed to make this connection. Only one model, Gemini 3.1 Pro, managed to get more than half of these answers correct. The others continued to insist that the character would look in the old, empty location, ignoring the fact that the character could see the toy.
The difficulty became even more pronounced when the researchers changed the prepositions used to describe the toy's location. In the standard story, a toy is placed "in" a box. In the test, the researchers described the toy as being placed "on" a box or "underneath" a bucket. When the story implied that the toy was visible because it was sitting on top of a container rather than hidden inside, every single model failed completely, scoring zero percent on these questions. They could not infer that the visibility of the object changed the character's state of mind. Similarly, when the story involved a character being explicitly told where the toy was moved, or when the question asked about what one character thought another character would do, the models stumbled. They often defaulted to simple, rigid patterns rather than dynamically tracking the mental states of the characters as the story unfolded.
The researchers also looked at how the models handled the gender of the characters. They found that the models performed slightly better when the two characters in the story shared the same gender, and slightly worse when they had different genders. This suggests that the models might be using gender as a shortcut to keep track of who is who, rather than truly understanding the roles and actions of each individual. When the shortcut was removed by using two characters of the same gender, the models seemed to rely more on the sequence of events, which paradoxically helped them get the answer right more often. This indicates that their reasoning is not as robust as human reasoning, which does not rely on such superficial cues to understand a situation.
The core finding of the study is that while these language models can mimic the correct answers to theory of mind tests, they do not seem to possess the underlying ability to understand mental states in the way humans do. They struggle to connect the dots between what a character can see, what they know, and what they believe. The models appear to be using "heuristics," or mental shortcuts, to guess the answer based on the most common patterns in their training data. When the story deviated from the standard pattern by adding visual details or changing the spatial relationships, the shortcuts failed. The researchers suggest that this is because the models are trained only on text. They have never seen a toy being moved, never experienced the feeling of looking for something, and never had to rely on their own eyes to update their knowledge of the world. Without this embodied experience, they cannot truly grasp the concept of a belief that is different from reality.
Ultimately, the study serves as a caution against assuming that because a machine can talk like a human, it thinks like one. The models can produce answers that look like they have a theory of mind, but when tested with scenarios that require genuine understanding of perception and belief, they reveal their limitations. They are excellent at repeating what they have read, but they are not yet capable of the flexible, common-sense reasoning that allows humans to navigate the complex social world. As these tools become more common, it is important to remember that their apparent intelligence is a reflection of the data they have consumed, not a sign that they have developed a mind of their own.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.