← Latest papers
🧬 biology

Using Theory of Mind to Arbitrate between Social and Non-social Learning

This paper proposes and validates a Rational Mentalizing model, which demonstrates that humans use Theory of Mind to rationally arbitrate between social and non-social learning by estimating the utility of observing others versus direct exploration.

Original authors: Lance Ying, Ryan Truong, Joshua B. Tenenbaum, Samuel J. Gershman

Published 2026-07-31
📖 4 min read☕ Coffee break read

Original authors: Lance Ying, Ryan Truong, Joshua B. Tenenbaum, Samuel J. Gershman

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). ⚕️ This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer

Imagine you are standing at a crossroads in a bustling city, trying to find the best taco truck. You have two choices: you can wander around yourself, checking every alleyway until you find one (which takes time and energy), or you can watch a stranger walking down the street. If you see them heading toward a specific corner, you might decide to follow them, assuming they know a good spot. But here's the catch: what if that stranger is actually looking for a pizza place, not tacos? If you follow them blindly, you'll end up hungry and confused. This everyday dilemma—deciding when to learn from others and when to figure things out on your own—is a central puzzle in the field of cognitive science. Scientists study how humans and animals gather information, but a key question remains: how do we know when to stop watching and start acting? To answer this, researchers rely on concepts like "Social Learning" (learning from others) and "Theory of Mind" (the ability to guess what someone else is thinking, like their goals or what they know). While we know we do these things, figuring out the exact mental math we use to balance the cost of watching against the benefit of learning has been a tricky challenge.

In this paper, researchers from Harvard and MIT propose a new way to understand this decision-making process, which they call the "Rational Mentalizing" model. Think of this model as a super-smart, internal calculator that doesn't just watch what people do, but tries to figure out why they are doing it. The authors built a computer simulation of a game where a player (you) controls a red character on a grid map. The goal is to find a treasure chest, but the path is blocked by colored barriers that can only be opened by finding a matching "amulet" held by one of several wizards. The twist? You don't know which wizard has the real amulet. You can either move your own character to check the wizards yourself (which costs time and effort) or you can pause and watch a non-player character (an NPC) move around. If you watch, the NPC moves one step while you stay still, giving you a clue about where the amulet might be.

The researchers ran four different versions of this game to see how people behave. In some versions, there was only one NPC and one goal; in others, there were multiple NPCs with different goals, or even a "novice" NPC who didn't know the map as well as an "expert." They found that humans are incredibly smart about this trade-off. We don't just copy everyone, nor do we ignore everyone. Instead, we subconsciously run a mental simulation: "If I watch this person for a few more steps, will I learn enough to save me from wandering around blindly?" If the answer is yes, we watch. If the answer is no (maybe the person is going to a different treasure chest than the one we want), we stop watching and start moving ourselves.

To prove this, the team compared real human players against three different computer models. The first model, the "Naive Observer," just watched everyone blindly. The second, the "Rational Observer," did a cost-benefit analysis but didn't try to understand the other person's thoughts. The third, the "Mentalizing Observer," tried to guess the other person's goals but didn't care about the cost of watching. None of these simple models could fully explain how humans behave. However, the "Rational Mentalizing" model, which combines guessing what the other person wants with calculating if it's worth the time to watch, matched human behavior almost perfectly. In fact, the model's predictions were so accurate that they were as close to human behavior as two different humans are to each other.

The study suggests that our brains are constantly running these complex calculations. We aren't just copying successful people; we are actively reasoning about their hidden goals and beliefs to decide if their actions will teach us something useful. If we realize someone is looking for something we don't care about, or if they are just as confused as we are, we stop watching and go our own way. The researchers found that when people ignored these mental calculations—either by watching too much without thinking about the cost, or by watching without understanding the other person's goals—they made worse decisions. This work suggests that the secret to our social intelligence isn't just in our ability to read minds, but in our ability to use that mind-reading to save our own time and energy. It turns out that being a good learner isn't just about paying attention; it's about knowing exactly when to look away.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →