← Latest papers
💬 NLP

LLMs and their Limited Theory of Mind: Evaluating Mental State Annotations in Situated Dialogue

This paper introduces a two-step framework using large language models to annotate shared mental models in team dialogues and detect discrepancies, revealing that while LLMs perform well on straightforward tasks, they systematically fail in scenarios requiring spatial reasoning or prosodic disambiguation.

Original authors: Katharine Kowalyshyn, Matthias Scheutz

Published 2026-06-30
📖 5 min read🧠 Deep dive

Original authors: Katharine Kowalyshyn, Matthias Scheutz

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are playing a high-stakes game of "Blindfolded Treasure Hunt" with a friend. You are the Searcher, walking through a real building with your eyes open. Your friend is the Director, sitting in a different room with a map that is slightly wrong. They can't see you; they can only talk to you over the phone.

To succeed, you both need to build a Shared Mental Model (SMM). This is like a invisible, mental whiteboard where you both write down what you know: "I'm in the hallway," "The box is blue," "You are looking for the green box." If your mental whiteboards don't match, you get lost, and the team fails.

This paper asks a simple but deep question: Can a super-smart AI (a Large Language Model or LLM) act like a referee and accurately track what's on your mental whiteboards just by listening to your phone conversation?

The Experiment: The "Blind" vs. The "All-Seeing"

The researchers set up a clever test using old recordings of humans playing this game. They created three groups of "referees" to analyze the same conversations:

  1. The Naive Humans: Regular people who only heard the audio. They had to guess what the Searcher was seeing based only on the words spoken.
  2. The AI Models: Three different AIs (o3-mini, Claude, and Gemma) that also only heard the audio.
  3. The "God's Eye" View (Ground Truth): A team of humans who watched the video of the game. They knew exactly where the Searcher was standing and what they were looking at at every second. This was the "correct answer key."

The Two-Step Process

The researchers used the AI in two different roles, like a writer and an editor:

  • Step 1: The Writer (Trace Generation): The AI listened to the conversation and tried to write down a "mental model trace." It had to guess: What does the Searcher believe right now? What is their goal? What have they promised to do?
  • Step 2: The Editor (Discrepancy Detection): A separate AI (acting as a judge) compared the "Writer's" guess against the "God's Eye" video answer key. It counted the mistakes.

The Mistakes: Where the AI Got Lost

The paper found that while the AIs were good at sounding smart, they struggled with the messy reality of human communication. Here are the main types of errors, explained with analogies:

  • The "Hallucinating" Director (Unsupported Beliefs):

    • What happened: The AI invented facts that weren't there.
    • Analogy: Imagine the Director says, "I think there's a cat in the room." The AI, hearing this, writes on the mental whiteboard, "The Searcher believes there is a cat." But the video shows there is no cat. The AI made up a belief just because the words sounded plausible.
    • Result: The AI (especially Claude) did this a lot, filling the whiteboard with imaginary details.
  • The "Missing" Details (Omissions):

    • What happened: The AI missed things that were actually said.
    • Analogy: The Searcher says, "I'm turning left." The AI's whiteboard stays blank, forgetting to update the location.
    • Result: The AI (especially o3-mini) was often too quiet, leaving out important updates.
  • The "Spatial" Confusion:

    • What happened: The AI got confused about where things were in space.
    • Analogy: If the Director says, "The box is to your left," the AI might write, "The box is to the right." It struggles to translate words into a 3D map in its head.
  • The "Stutter" Trap:

    • What happened: Humans say "um," "uh," and stumble over words. The AI often treated these stutters as new information.
    • Analogy: If the Director stutters, "There's a... uh... door," the AI might think, "Aha! The 'uh' means they are changing their mind about the door!" and update the whiteboard unnecessarily.

The Verdict: Smart but Groundless

The paper concludes that current AIs are like actors who are very good at reading a script but terrible at improvising in a real room.

  • They can mimic the language of thinking: They can use words like "I believe" or "My goal is."
  • They cannot ground the thinking: They can't connect those words to the physical reality (the video) or handle the messy, stuttering nature of real human speech.

Interestingly, the "Naive Humans" (who also didn't see the video) made similar mistakes to the AIs. This proves that the task is genuinely hard without seeing the video. However, the AIs were worse at avoiding "hallucinations" (making things up) than the humans were.

The Bottom Line

This study doesn't say AIs can't be used in teams. It says that if you want an AI to understand a team's shared mental state just by listening to them talk, it currently lacks the "common sense" to know what is real and what is just a guess. It needs the "video" (the real-world context) to truly understand what's happening, rather than just guessing based on the words it hears.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →