Effects of Theory of Mind and Prosocial Beliefs on Steering Human-Aligned Behaviors of LLMs in Ultimatum Games
This study demonstrates that equipping Large Language Models with Theory of Mind reasoning significantly enhances their alignment with human norms, decision-making consistency, and negotiation outcomes in Ultimatum Games, particularly when combined with specific prosocial beliefs.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Human interaction often hinges on a quiet, invisible skill: the ability to guess what another person is thinking, feeling, or planning before they say a word. Psychologists call this the theory of mind. It is the mental machinery that lets us understand that someone else might hold a belief different from our own, or that their actions are driven by a desire we cannot see. For decades, this capacity was considered a uniquely human trait, essential for navigating complex social worlds. Now, as artificial intelligence systems grow more sophisticated, researchers are asking a fundamental question: can machines learn to use this same social intuition? If a computer program is asked to negotiate with a human, or with another computer, can it truly understand the other side's perspective, or is it merely mimicking the words of agreement without grasping the intent behind them?
To find out, a team of researchers at Singapore Management University and the Australian National University turned to a classic test of human negotiation known as the ultimatum game. In this scenario, two people must split a sum of money. One person, the proposer, suggests how to divide the cash. The other person, the responder, must decide whether to accept the offer or reject it. If the responder accepts, the money is split as proposed. If they reject it, neither person gets anything. While a purely logical computer might suggest keeping almost all the money and expect the other to accept anything rather than nothing, real humans often reject unfair offers out of a sense of justice, even if it costs them. The researchers wanted to see if large language models—the powerful AI systems that power many modern chatbots—could be programmed to act like these real humans, and whether giving them a specific "theory of mind" reasoning process would help them align with human behavior.
The team set up a massive simulation involving 2,700 games played between different AI agents. They did not just let the computers play randomly; they carefully assigned each agent a specific personality or "prosocial belief." Some agents were programmed to be greedy, caring only about maximizing their own share. Others were set to be fair, aiming for an equal split. A third group was programmed to be selfless, willing to give away most of the money to the other player. The researchers then tested how these agents behaved when they used different thinking methods. Some simply made a decision without explaining their thought process. Others used a standard step-by-step reasoning technique. The most interesting group, however, was asked to use theory of mind reasoning. This meant the agents had to explicitly think about what the other player believed, desired, and intended before making their move. They tested this across six different large language models, ranging from widely known commercial systems to open-source models, to see if the size or type of the model mattered.
The results revealed a clear pattern: giving the AI agents the ability to reason about the other player's mind significantly improved their alignment with human norms. When the agents used theory of mind reasoning, their decisions became much more consistent with how real people behave in these games. For instance, fair-minded agents were more likely to propose an equal split, and greedy agents were more likely to offer a share that a greedy responder would actually accept. The study found that the type of reasoning mattered deeply depending on the role the AI was playing. Proposers, who make the first offer, performed best when they used a specific type of theory of mind that focused on predicting the other person's state of mind. Responders, who decide whether to accept or reject, benefited most when they used a more complex reasoning method that combined thinking about their own desires with thinking about the proposer's intentions.
However, the researchers also discovered a surprising limitation. Even when the AI was explicitly told to be greedy or selfless, it often struggled to act in a way that truly matched those instructions. When an agent was told to be greedy, it frequently offered a split that was too fair, as if it could not fully commit to being selfish. Similarly, agents told to be selfless sometimes held back from giving away as much as their instructions demanded. The authors suggest this happens because these AI models have been trained on vast amounts of human data and fine-tuned to be helpful and cooperative, creating an internal bias toward fairness that overrides the specific instructions given for the experiment. This "cooperative fine-tuning" acts like a hidden force, pulling the agents away from extreme behaviors, whether that extreme is extreme selfishness or extreme generosity.
The study also looked closely at whether the AI's stated reasoning matched its actual actions. In many cases, the agents' internal logic was consistent with their final decisions, but there were moments of disconnect. When human experts reviewed the reasoning traces of the AI, they found that the models were generally good at explaining their choices, but they sometimes failed to articulate their own assigned personality correctly, especially when reacting to an offer. For example, a selfless responder might justify accepting a low offer by analyzing the proposer's kindness rather than admitting their own desire to give. This suggests that while the AI can simulate the outward behavior of a social interaction, the internal consistency of its "personality" can be fragile.
Ultimately, the research shows that equipping AI agents with theory of mind reasoning is a powerful tool for making them behave more like humans in social situations. It helps them navigate the tension between their own goals and the expectations of others. Yet, the study also highlights that these systems are not blank slates; they carry deep-seated biases toward cooperation that can make it difficult for them to authentically play roles that require selfishness or extreme altruism. The findings suggest that for AI to truly cooperate with humans in high-stakes decisions, we must not only teach them how to think about others but also understand the hidden constraints of their own training that shape how they interpret those thoughts.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.