← Latest papers
🤖 AI

What Do You Think I Think? Accounting for Human Beliefs Using Second-Order Theory of Mind

This paper presents a second-order Theory of Mind framework using I-POMDPs that enables an agent to detect and adaptively address human cognitive biases and erroneous beliefs about the agent, thereby improving interaction quality and feedback utility as demonstrated in a user study.

Original authors: Patrick Callaghan, Reid Simmons, Henny Admoni

Published 2026-05-14
📖 4 min read☕ Coffee break read

Original authors: Patrick Callaghan, Reid Simmons, Henny Admoni

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are teaching a robot how to sort a deck of special cards. You have a secret rule in your head (e.g., "Red cards go in Bin 1, Blue cards go in Bin 2"). The robot doesn't know the rule yet; it has to learn it by watching you sort the cards and listening to what you say.

The problem is that you might be making mistakes in how you think the robot thinks.

The Core Problem: The "Mind-Reading" Gap

In this paper, the researchers are trying to solve a specific communication breakdown. Humans often have "cognitive biases"—mental shortcuts that make us assume things about others that aren't true.

  • The Scenario: You might think, "The robot is smart; it must know that Red means Bin 1 because that's the most obvious pattern."
  • The Reality: The robot is actually confused and thinks, "I have no idea what Red means yet."
  • The Result: You stop teaching because you think the robot knows, or you teach in a way that confuses the robot because you're assuming it understands your shortcuts.

The Solution: The "Second-Order" Mind

The researchers built a robot with a special mental superpower called Second-Order Theory of Mind (ToM-2).

To understand this, think of it like a game of "I know that you know that I know":

  1. Zero-Order Robot (The Basic Bot): It only looks at the cards you put in the bins. It thinks, "Okay, you put a Red card here. I will update my rule." It has no idea what you are thinking.
  2. Second-Order Robot (The Mind-Reader): This robot doesn't just look at the cards; it simulates a model of your brain. It asks itself: "What does the human think I know right now?"

If the robot realizes, "Oh, the human thinks I know the rule is about Color, but I actually think it's about Shape," it can spot the disconnect.

How the Robot Fixes the Conversation

The paper describes a game where the robot has to guess a sorting rule. The robot uses its "mind-reading" ability to detect two specific human habits (biases):

  1. Representativeness Heuristic: You might focus only on the cards that look like the rule you have in mind and ignore the ones that prove other rules wrong.
  2. Confirmation Bias: You might only pay attention to the robot's feedback if it confirms what you already believe.

When the robot detects these habits, it changes its strategy. Instead of just saying, "I'm not sure," it gives you a two-part nudge:

  • Part A (The Nudge): "It seems like you want me to think the rule is about Shape, but..."
  • Part B (The Correction): "...it's still possible the rule is about Color."

By explicitly stating, "I know what you think I'm thinking, but I'm actually thinking something else," the robot breaks the human's mental loop and forces them to reconsider their teaching strategy.

The Experiment: Teaching Cards

The researchers tested this with 30 real people.

  • The Setup: People used a computer to teach a virtual agent to sort cards with different colors, shapes, and numbers.
  • The Comparison:
    • Group A taught a "Basic Bot" (Zero-Order) that couldn't read minds.
    • Group B taught the "Mind-Reading Bot" (Second-Order) that could detect biases.

What Happened?

  1. Better Teaching: People taught the Mind-Reading Bot much more efficiently. Because the bot corrected their misunderstandings, the humans stopped making redundant moves and gave better examples.
  2. More Information: The cards people chose to show the Mind-Reading Bot contained more useful information. The humans were subconsciously adjusting their teaching style because the bot was "talking back" in a way that made sense of their confusion.
  3. The "Unconscious" Benefit: Interestingly, when asked afterwards, the humans didn't say, "Wow, that bot was amazing at reading my mind!" They just felt the feedback was more useful. They didn't necessarily realize why the interaction was smoother; they just knew it worked better.

The Bottom Line

This paper shows that if an AI can model what you think it knows (and realize when you are wrong about that), it can give you better feedback. This helps you teach it faster and more accurately, even if you don't fully realize the AI is doing the heavy lifting to fix your mental blind spots.

Note: The researchers tested this in a simple card-sorting game. They admit that real life is much more complex, and their robot only looked at two specific types of human biases, but the core idea—that modeling human misconceptions improves teamwork—worked in their experiment.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →