Grounded-Exo2Ego: Structured Semantic Grounding for Robust Exocentric-to-Egocentric Video Generation
This paper introduces Grounded-Exo2Ego, a dual-branch video diffusion framework that combines geometric anchoring with a novel semantic grounding branch and a camera re-localization algorithm to robustly generate egocentric videos from exocentric inputs, achieving state-of-the-art performance on the EgoExo4D dataset through improved architecture and fully automated synthetic data generation.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine watching a video of a person cooking from across the kitchen. You see the whole scene: the stove, the counter, the chef's movements. Now, imagine being able to instantly switch that perspective to see exactly what the chef sees, as if you were wearing their eyes. This ability to transform a third-person view into a first-person one is the goal of a new scientific effort called exocentric-to-egocentric video generation. It is a crucial step for building robots that can learn by watching human videos and for creating immersive virtual reality experiences where users can relive moments from another person's perspective. However, turning a video shot from the outside into a view from the inside is incredibly difficult. Standard computer vision tools, which work well for changing angles in a static scene, often fail here. When the camera moves from a wide view to a close-up, the computer struggles to guess what is hidden behind objects or how the scene should look from a completely different angle, often resulting in distorted, blurry, or missing parts of the image.
A team of researchers at NVIDIA has developed a new system called Grounded-Exo2Ego to solve this problem. Instead of relying on the usual methods that try to force a geometric fit between the outside view and the inside view, they built a framework that understands both the shape of the room and the specific objects inside it. The researchers found that previous attempts failed because the computer's best guess at the 3D shape of the room was often wrong, especially for parts of the scene that were hidden from the original camera. When the computer tried to use these flawed 3D maps to generate a new view, the result was a video full of holes and warped surfaces. To fix this, the new system uses a dual-branch approach. One branch acts as a structural guide, using a rough 3D map of the scene to tell the computer where walls and floors should be. The second branch, which is the key innovation, acts as a semantic guide. It identifies specific objects in the video, such as a person, a bicycle, or a cooking pot, and tells the computer exactly what those objects should look like and where they belong, even if the 3D map is incomplete or broken.
The researchers also discovered a hidden flaw in how these systems were being trained. In the real world, the camera a person wears does not perfectly match the computer's 3D reconstruction of the room. This mismatch meant that when the computer tried to learn from real video data, it was essentially trying to learn from a distorted map that did not line up with the reality it was supposed to mimic. To solve this, the team created a new process that re-aligns the camera's position within the computer's imperfect 3D map before training begins. This ensures the computer learns from a consistent picture. Furthermore, to provide the system with perfect examples to learn from, they built a fully automated engine that creates thousands of synthetic videos. In these computer-generated worlds, they place animated 3D characters in various rooms and record them from both outside and inside perspectives, providing the model with flawless data that real-world cameras cannot easily capture.
When tested on a challenging dataset of real-world videos featuring activities like sports, cooking, and bike repair, the new system significantly outperformed existing methods. It produced videos that were sharper, more accurate, and much better at keeping objects in the correct place. For instance, while older methods often failed to reconstruct people entirely or placed them in the wrong spots, the new system successfully generated clear views of the subjects and their surroundings. The results showed that by combining a rough understanding of the room's layout with a detailed understanding of the objects within it, and by training on data that was carefully aligned and supplemented with synthetic examples, the computer could finally bridge the gap between watching a scene and seeing through someone else's eyes. This work suggests that for machines to truly understand and replicate human experiences, they need to look beyond simple geometry and learn the meaning of the objects they are trying to render.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.