← Latest papers
🤖 AI

Towards Embodied Cognition in Robots via Spatially Grounded Synthetic Worlds

This paper introduces a publicly available synthetic dataset generated in NVIDIA Omniverse to train Vision-Language Models on spatial reasoning and Visual Perspective Taking, serving as a foundational step toward embodied cognition in human-robot interaction.

Original authors: Joel Currie, Gioele Migno, Enrico Piacenti, Maria Elena Giannaccini, Patric Bach, Davide De Tommaso, Agnieszka Wykowska

Published 2026-07-29
📖 4 min read☕ Coffee break read

Original authors: Joel Currie, Gioele Migno, Enrico Piacenti, Maria Elena Giannaccini, Patric Bach, Davide De Tommaso, Agnieszka Wykowska

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are playing a game of hide-and-seek with a friend in a giant, cluttered living room. You are hiding behind the sofa, and your friend is standing by the door. If you shout, "I see the red ball to my left!", your friend has to do something surprisingly tricky: they can't just look left from their own spot. They have to mentally step into your shoes, spin around to see the room from your angle, and then figure out where that red ball is relative to them. This mental gymnastics is called "Visual Perspective Taking." It's a superpower humans have that lets us understand what others see, think, and feel. For robots to truly hang out with us, play with us, or help us in our homes, they need this same superpower. Right now, robots are often like people who can only see the world from their own eyes; they struggle to understand that "left" for them might be "right" for you. This paper dives into the world of Artificial Intelligence (AI) to see if we can teach robots to do this mental spinning act, specifically by using a special kind of training data that acts like a video game simulator.

The researchers behind this paper, Joel Currie and his team, are tackling a specific problem: robots and the AI brains that drive them are great at recognizing objects, but they are terrible at understanding exactly where those objects are in 3D space, especially when looking at them from different angles. While some older robot methods use rigid math rules to figure out space, the team suggests that modern "Vision-Language Models" (AI that can see pictures and read text) might be the future, but only if we can teach them better. The paper argues that these smart models are currently failing at spatial reasoning not because they are "dumb," but because they haven't been trained on enough data that perfectly links what an object looks like to exactly where it is.

To fix this, the team created a brand-new "training gym" for robots, but instead of a real room, they built a digital one inside a powerful video game engine called NVIDIA Omniverse. Think of it as a digital sandbox where they can drop in a single, slightly squished cube, take a picture of it from a camera with a random height, and write a sentence describing it. The magic part is that because they built this world in a computer, they know the exact math of where the cube is, down to the millimeter. They generated a dataset full of these scenes, pairing an image, a language prompt, and a precise "4×4 transformation matrix" (a fancy math grid that tells you exactly how the object is positioned).

In this first experiment, the team didn't ask the AI to solve the whole puzzle of 3D space all at once. Instead, they focused on a simpler, foundational skill: guessing how far away the object is along the Z-axis (the depth line relative to the camera), while keeping everything else still. They want to see if the AI can look at the picture and the description and correctly guess that distance. If the AI can learn this in their perfect, synthetic world, the hope is that it will eventually learn to handle the full, messy complexity of a real room. The paper presents this as a "proof-of-concept," meaning it's an early-stage test to see if the idea works. They aren't claiming to have solved robot intelligence today; rather, they are laying the first brick of a foundation. They have made this dataset public so other scientists can use it to train their own robots, aiming for a future where robots can truly understand "what others see" and navigate our world with the same spatial awareness we take for granted.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →