Beyond Sally-Anne: Evaluating Theory of Mind in LLMs using Epistemic Schelling Points
This paper introduces the Epistemic Asymmetry Schelling Task (EAST), a novel dialogue-based benchmark that reveals significant gaps in the functional Theory of Mind and epistemic tracking abilities of Large Language Models, demonstrating that even frontier models struggle with dynamic social coordination despite high performance on traditional static tests.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you and a friend are playing a high-stakes game of "Secret Word" over text messages. You can't talk to each other, but you both have to pick the exact same word from a list of four to win. The twist? You both have secret "personas" (like being a tailor or a surgeon), and you need to guess what word the other person will pick based on what they know about you.
This is the setup for a new study by researchers at Google, who wanted to see if Large Language Models (LLMs) – the super-smart AI brains behind chatbots – can actually "think" about what other people are thinking. This ability is called Theory of Mind.
The Old Way: The "Sally-Anne" Trap
For a long time, scientists tested AI on Theory of Mind using a classic riddle called the "Sally-Anne task." It's like a storybook quiz: "Sally puts a ball in a box and leaves. Anne moves it to a basket. Where will Sally look?"
The researchers argue that this old way is a bit like a video game boss that players have memorized. Because AI models are trained on massive amounts of text from the internet, they've likely seen this exact riddle a million times. They might just be guessing the answer based on the story's pattern, not because they truly understand that Sally has a different perspective than Anne. It's like a parrot repeating a phrase it heard in a movie; it sounds smart, but it doesn't know what the words mean.
The New Game: EAST (The "Mind-Reading" Coordination Game)
To fix this, the team invented a new game called EAST (Epistemic Asymmetry Schelling Task). Instead of a storybook, it's a live coordination challenge.
Here's how it works:
- The Setup: Two AI agents are paired up. Each gets a secret job (like "Bespoke Tailor" or "General Surgeon").
- The List: They are shown four words. Two words are specific to their own jobs (e.g., "fabric" for the tailor, "scalpel" for the surgeon). One word is a mix of both (e.g., "stitches"). One word is random (e.g., "bicycle").
- The Goal: They must pick the same word without talking.
- The Twist: The game changes the rules of "who knows what."
- Symmetric: Both know each other's jobs.
- Asymmetric: One knows the other's job, but the other doesn't know the first one's.
- Zero Knowledge: Neither knows the other's job.
To win, the AI has to do something tricky: it has to suppress its own "favorite" word and guess what the other person is thinking, considering what that person knows (or doesn't know) about the AI.
What the Game Revealed
When the researchers ran this game with 14 different AI models (ranging from tiny 1-billion-parameter models to the massive, state-of-the-art "frontier" models), the results were a bit of a reality check.
1. The "Big" Models are the Only Ones Who Got It
Only the very largest, most advanced models (like the frontier-class Gemini 3.1 Pro) managed to win consistently. They could successfully navigate the tricky "Asymmetric" and "Zero Knowledge" rounds. The smaller, less powerful models? They mostly failed. They treated the game like a simple word association test, picking their own job-related word and hoping for the best.
2. The "Mind-Reading" Glitch
The researchers looked inside the AI's "brain" (its reasoning steps) to see why it failed. They found a specific type of error: Epistemic Tracking Errors.
Imagine you are trying to guess what your friend will eat for lunch. You know they love pizza. But you also know they don't know that you love pizza. If you pick "pizza" because you love it, you've made a mistake. You forgot that your friend doesn't know your taste.
The AI models did exactly this. They often confused their own private knowledge with what the other player knew.
- In the Asymmetric round, the AI that knew the other player's job failed to realize the other player was "blind" to its own job. It picked a word based on its own secret identity, thinking the other player would guess it too.
- In the Zero Knowledge round, the models couldn't handle the "void." They couldn't imagine a world where neither of them had any clues, so they just grabbed their own favorite word.
3. The "Thinking Aloud" Trick
The researchers tried helping the AI by asking it to "think step-by-step" (a method called Chain-of-Thought). For the smartest models, this helped them stop guessing randomly and start using real logic. But for the weaker models, just asking them to think didn't fix the problem; they still got confused about who knew what.
The Bottom Line
The paper suggests that while AI models are getting really good at answering questions on tests, they are still struggling with robust social reasoning. They can mimic the words of a social interaction, but they often fail to truly track the mental states of others, especially when the rules of "who knows what" get complicated.
The authors aren't saying AI is "broken" or that it will never understand people. They are saying that right now, the ability to keep a clear mental boundary between "what I know" and "what you know" is a major bottleneck. Until AI can reliably stop projecting its own secrets onto others, it might not be ready for the messy, unpredictable world of real human conversation.
In short: The AI is smart enough to play the game, but it's still learning how to stop thinking the whole world is inside its own head.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.