← Latest papers
💬 NLP

Contextual Observer Grounding: Evaluating Situated Spatial Reasoning in Vision-Language Models

This paper introduces the Point-of-View Benchmark (POVBench) to evaluate the ability of vision-language models to perform "contextual observer grounding"—inferring a speaker's situated perspective from contextual cues—and demonstrates that while current models struggle with this task, explicit breakdowns of observer-relative spatial reasoning can significantly improve target localization.

Original authors: Mimo Shirasaka, Haochen Zhang, Yonatan Bisk

Published 2026-09-09
📖 5 min read🧠 Deep dive

Original authors: Mimo Shirasaka, Haochen Zhang, Yonatan Bisk

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine trying to find a lost item in a room you have never entered, based only on a friend's description of where they left it. You might hear, "I dropped my keys on the table to the left of the lamp while I was making coffee." To solve this puzzle, you do not just need to know what a table and a lamp look like; you must first reconstruct your friend's position in the room, imagine their line of sight, and then mentally place the keys relative to that specific viewpoint. This ability to understand space from another person's perspective, using clues about their activity and the environment, is a fundamental part of human communication. It allows us to collaborate efficiently without needing to share every single visual detail. For robots and artificial intelligence to work alongside us in our homes, they must master this same skill, known as situated spatial reasoning.

A new study by researchers at the University of Tokyo and Carnegie Mellon University investigates whether modern artificial intelligence systems can truly perform this task. The team, led by Mimo Shirasaka, Haochen Zhang, and Yonatan Bisk, created a rigorous test called the Point-of-View Benchmark to see if machines can infer a speaker's location and then locate an object based on that inferred perspective. Their work reveals that while current AI models are impressive at many visual tasks, they struggle significantly when asked to reason about space from a viewpoint they cannot directly see, even when given clear instructions.

The researchers built their test using a computer-generated world of procedurally created houses, complete with realistic furniture and layouts. In this digital environment, a simulated robot explores each house, capturing a series of first-person images as it moves from room to room. The core of the experiment involves a natural language sentence describing an object's location, such as "I saw the book to the left of the bed." The challenge for the artificial intelligence is to look at the robot's exploration photos, figure out where the person speaking was standing when they said that sentence, and then point to where the book would be in that specific view. The study tested three different levels of difficulty to understand exactly where the AI gets stuck. In the hardest version, the sentence only gives a hint about the activity, like "While replacing a light bulb," forcing the AI to guess the location. In a medium version, the sentence explicitly names the object the person was interacting with, such as "While replacing the light bulb on the floor lamp." In the easiest version, the AI is simply shown the exact photo from the speaker's perspective.

The results were revealing. Across a wide range of advanced AI models, including both commercial systems and open-source versions, performance remained low in all three scenarios. Even when the AI was explicitly told which object the speaker was near, or when it was given the exact photo from the speaker's viewpoint, the models frequently failed to pinpoint the correct location. The gap between the hardest and medium versions was surprisingly small, suggesting that the main difficulty is not just figuring out where the speaker was standing, but rather understanding how to translate a description like "to the left of" into a specific spot in a visual scene. The models often picked the wrong room or confused the reference object, such as mistaking a door frame for a refrigerator, because they could not effectively combine the visual clues with the context of the activity.

To understand if the models could improve with help, the researchers tried a different approach. They asked the AI to break down its thinking process step by step, explicitly stating how it identified the speaker's position, located the reference object, and then determined the direction of the target. This method, which the authors call structured spatial reasoning, led to a noticeable improvement in accuracy. When the models were guided to reason through the problem in this structured way, they became better at locating the hidden objects. This suggests that the AI possesses the necessary pieces of information but struggles to assemble them into a coherent mental model without explicit guidance on how to connect the dots.

The study also tested these ideas in the real world using actual video footage of people walking through houses. The results mirrored the computer simulations: the AI models continued to struggle with the task, performing at or below a fifty percent success rate. However, the models that used the step-by-step reasoning approach again showed better results than those that did not. This indicates that the difficulty is not just a flaw of the computer-generated environment but a genuine challenge for current technology when faced with the messy, complex reality of human spaces.

Ultimately, this research highlights a critical gap in the capabilities of embodied artificial intelligence. While machines can recognize objects and answer questions about what they see, they have not yet learned to naturally infer the perspective of a human speaker or to reconstruct a scene from a viewpoint they have never directly observed. The findings suggest that for robots to truly collaborate with humans in our daily lives, they need to develop a deeper ability to reason about space from another person's point of view, using context and common sense to fill in the missing pieces of the visual puzzle.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →