← Latest papers
💻 computer science

Communicating about Space: Language-Mediated Spatial Integration Across Partial Views

This paper introduces COSMIC, a benchmark for evaluating how Multimodal Large Language Models (MLLMs) collaborate via dialogue to build shared spatial understanding from partial views, revealing that while models can identify shared objects, they significantly lag behind humans in relational reasoning and maintaining globally consistent mental models of the environment.

Original authors: Ankur Sikarwar, Debangan Mishra, Sudarshan Nikhil, Ponnurangam Kumaraguru, Aishwarya Agrawal

Published 2026-03-31
📖 5 min read🧠 Deep dive

Original authors: Ankur Sikarwar, Debangan Mishra, Sudarshan Nikhil, Ponnurangam Kumaraguru, Aishwarya Agrawal

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you and a friend are trying to meet up in a massive, unfamiliar park, but neither of you has a map. You are standing on opposite sides of a large fountain.

  • You can see a tall oak tree and a red bench, but the fountain is hidden behind a hill.
  • Your friend can see the fountain and a blue slide, but the oak tree is blocked by a fence.

To find each other, you can't just look at your own view. You have to talk. You say, "I see a red bench near a tree." Your friend replies, "I see a fountain near a blue slide." Slowly, by combining your descriptions, you build a shared mental picture of the whole park in your heads. You realize, "Ah! The red bench is right next to the fountain!"

This paper is about teaching AI robots (specifically, advanced "Multimodal Large Language Models" or MLLMs) to do exactly this. The researchers created a test called COSMIC to see if AI can learn to build that shared mental picture just by chatting with each other.

Here is the breakdown of what they found, using some fun analogies:

1. The Test: "The Blindfolded Roommates"

The researchers set up a digital experiment where two AI agents are placed in a 3D room.

  • Agent A (The Answerer) has a question but can only see half the room.
  • Agent B (The Helper) has the rest of the view but doesn't know the question.
  • The Goal: They must chat back and forth to solve the puzzle.

The test had five levels of difficulty, like climbing a ladder:

  1. Spotting the Common Ground: "Do you see the lamp I'm looking at?" (Easy)
  2. Counting Together: "How many chairs are in the whole room?" (Harder, because you have to make sure you aren't counting the same chair twice).
  3. Measuring Distances: "Which object is closer to the lamp?" (Requires combining views to measure).
  4. Changing Perspectives: "From my point of view, where is the door?" (The Helper sees the door; the Answerer has to imagine it from their own angle).
  5. Drawing the Map: "Is this top-down map of the room correct?" (The hardest level: building a complete 3D map in your head from two 2D pictures).

2. The Results: The AI is "Chatty but Clueless"

The researchers compared the AI's performance to real humans. The results were a mix of "not bad" and "disastrous."

  • The Good News: The AI is pretty good at the easy stuff. If you ask, "Do you see the red chair?", the AI can usually say "Yes." It's like a tourist who can point out landmarks.
  • The Bad News: As soon as the task gets tricky (like figuring out where things are relative to each other or drawing a map), the AI falls apart.
    • Humans: Got 95% of the answers right. They were efficient, using few words to lock onto the right objects and quickly building a shared map.
    • The Best AI: Got only 72% right.
    • The Worst AI: Was barely better than guessing randomly on the hardest tasks.

3. The Problem: The AI is Like a Tourist with a Broken Compass

The paper found three main reasons why the AI fails:

  • The "Hallucination" Trap: Sometimes the AI sees a chair and calls it a "sofa." Once it makes that mistake, the whole conversation goes off the rails. It's like trying to navigate a city when your GPS keeps telling you the bakery is a gas station.
  • The "Double-Counting" Confusion: When counting objects, the AI often gets confused about whether the chair it sees is the same chair its partner sees. It ends up counting the same chair twice, or missing one entirely.
  • The "Perspective" Failure: This is the biggest issue. Humans are great at saying, "If I'm facing North, and you see the door to your Left, that means the door is behind me." The AI struggles to rotate the world in its head. It gets stuck in its own viewpoint and can't imagine what the other person sees.

4. The "Thinking" Feature Didn't Help Much

The researchers tried turning on the AI's "Thinking Mode" (where the AI pauses to reason before speaking).

  • Did it help? Yes, for simple things like spotting objects.
  • Did it fix the hard stuff? No. Even when the AI "thought" hard, it still couldn't build a consistent map of the room. It's like a calculator that is very good at adding numbers but still can't understand the concept of a map.

5. The Verdict: We Have a Long Way to Go

The study concludes that while AI is getting smarter at seeing and talking, it is still terrible at collaborating to build a shared reality.

  • Humans are like expert detectives: They quickly agree on the facts, stop guessing, and solve the case.
  • AI is like a nervous tourist: It talks a lot, keeps suggesting new possibilities, gets confused about what the other person is seeing, and rarely settles on a single, correct picture of the world.

In short: If you want an AI to help you navigate a new building or coordinate a team of robots, you can't just ask it to "chat." Right now, the AI is too likely to get lost in its own head and fail to build the shared map needed to get the job done. The researchers hope this test (COSMIC) will help engineers fix these specific blind spots so AI can truly work with us, not just near us.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →