← Latest papers
🤖 AI

Belief Without Behavior: Measuring the Translation of Theory of Mind into Coordinated Social Action in Vision-Language Models

This paper introduces MOSAIC, a multimodal benchmark revealing that current Vision-Language Models fail to translate Theory of Mind reasoning into coordinated social actions due to bottlenecks in generating coherent nonverbal signals and interpreting others' behaviors, whereas models with explicit belief-action coupling succeed.

Original authors: Tonglin Yan, Gregoire Sergeant-Perthuis, David Rudrauf

Published 2026-08-24
📖 5 min read🧠 Deep dive

Original authors: Tonglin Yan, Gregoire Sergeant-Perthuis, David Rudrauf

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Human connection relies on a silent, rapid exchange that goes far beyond words. When we speak, we also shift our gaze, adjust our posture, and shift our facial expressions, all while trying to guess what the other person is thinking and feeling. This ability to understand another person's mind and then use that understanding to guide our own actions is known as the theory of mind. For decades, scientists have studied this capacity, often by asking people to answer questions about stories or by watching how they react to simple social puzzles. The goal has been to see if artificial intelligence can do the same. But a critical gap has remained: knowing what someone thinks is different from acting on that knowledge in real time. A system might correctly guess that a person is hiding a secret, but can it then decide to look away, smile, or move its body in a way that subtly encourages that person to reveal it? Until now, no test has been able to measure whether an AI can translate a mental guess into a coordinated, physical social signal.

Researchers at Université Paris-Saclay and Sorbonne Université set out to close this gap with a new experiment called MOSAIC. They built a virtual world where two digital agents interacted in a game of treasure hunting. In this scenario, one agent, the subject, knew where a hidden reward was located, while the other, the participant, did not. The subject's task was to either help the participant find the treasure or, depending on the rules of the round, trick them into choosing the wrong box. The game was designed to be a rigorous test of social intelligence. Over ten rounds of interaction, the agents had to communicate using a mix of spoken words, body movement, eye direction, and facial expressions. The researchers varied the difficulty by changing the rules: sometimes the subject was supposed to be honest, and other times it was supposed to be deceptive. Crucially, they also changed the subject's internal instructions, asking some to act on their own desires without thinking about the other person, while asking others to explicitly model the participant's beliefs and try to influence them.

The researchers tested thirteen different artificial intelligence models in this environment, including eleven vision-language models that can see and speak, and one text-only model. They ran two hundred trials for each model to gather enough data to see clear patterns. The results were stark. Most of the advanced AI systems failed to produce the coordinated behavior required to succeed. When asked to deceive, the models could not generate the necessary mix of misleading words and contradictory body language. When asked to help, they often failed to move or look in a way that guided the other agent. The study identified two specific breakdown points in the process. First, the majority of the models could not produce nonverbal signals that made sense directionally; their movements and gaze were often random or pointed in the wrong direction. Second, even when a model did produce a clear signal, the other agent in the simulation failed to notice it or act on it. The AI agents seemed to be talking and moving, but they were not truly communicating.

One model stood out as a complete exception. A hybrid system called PCM-LLM, which was built with a specific, structured module designed to handle beliefs and intentions, succeeded in every condition. It could successfully guide the participant to the treasure when asked to help, and it could successfully mislead them when asked to deceive. This success suggests that the failure of the other models was not due to the difficulty of the task itself, but rather to a missing piece in their architecture. The standard models, which process language and images together, appeared to lack the internal mechanism to link a belief about another person directly to a physical action. The researchers found that for most models, the visual input they received actually acted as a distraction rather than a helpful clue. In some cases, removing the visual image entirely improved the model's ability to send a clear directional signal, suggesting that the visual data was interfering with their reasoning rather than supporting it.

The study also explored whether these models could learn to do better with practice. The researchers took one of the failing models and trained it using examples of successful behavior generated by the high-performing hybrid system. While this training changed the model's behavior slightly, it did not fix the fundamental problem. The model still could not coordinate its words, eyes, and body into a single, coherent social signal. This suggests that simply showing an AI how to move is not enough; it needs a deeper architectural change to understand that its actions are meant to influence another mind. The findings indicate that current AI systems, despite their ability to discuss complex social scenarios in text, have not yet learned to embody those ideas. They can talk about theory of mind, but they cannot yet live it. The gap between inferring what someone thinks and acting on that inference in a coordinated, physical way remains a distinct and measurable limitation of current technology.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →