XP-JEPA: Cross-Predictive Physics Grounding for Forecastable Latent Dynamics
The paper introduces XP-JEPA, a cross-predictive learning framework that grounds visual latent dynamics in privileged physical trajectories during training to significantly improve forecastability and control performance, while discarding the physical branch for visual-only deployment.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Robots that can plan ahead often rely on a mental shortcut: they build an internal map of the world, a kind of simulation where they can test out different moves before actually making them. To do this, the robot takes a picture of its surroundings, turns that image into a simplified set of numbers, and then asks a computer program to guess what those numbers will look like a moment later if the robot pushes, lifts, or turns something. If the guess matches the goal, the robot executes the move. This process works well when the robot's internal map accurately reflects how the physical world changes. However, a common problem arises when the robot learns to predict the future based only on pictures. It can become too clever for its own good, finding patterns in the visual data that are easy to predict but have nothing to do with real physics. For instance, the robot might learn that a certain shade of gray usually follows a certain shade of blue, even if the actual object hasn't moved at all. This creates a fragile understanding where the robot's predictions drift away from reality, causing it to fail when it tries to interact with the world.
A team of researchers at the National University of Singapore has developed a new method to fix this flaw, ensuring that the robot's internal predictions stay grounded in the laws of physics. They call their approach XP-JEPA. Instead of letting the robot learn solely from watching video, they gave it a second stream of information during training: a direct readout of the physical state of the objects, such as their exact position, size, and orientation in space. This physical data is like a hidden layer of truth that the robot cannot see with its camera but can access while it is learning. The researchers designed a system where the robot's visual brain and its physical brain work together. Both sides receive the same history of actions and try to predict the future. The visual side predicts what the next picture will look like, while the physical side predicts the next set of coordinates. Crucially, the system forces these two predictions to agree with each other. If the visual side predicts that a block will move up, but the physical side predicts it will stay still, the system learns that the visual prediction is wrong. This cross-checking process teaches the visual system to pay attention to the real physical changes that matter, rather than just guessing based on superficial visual patterns.
The researchers tested this method on a wide variety of robotic tasks, including pushing blocks, stacking objects, and tossing items into containers. They compared their new system against standard models that learn only from pictures. The results were striking. When the researchers let the new system run a simulation of a task, the error in its prediction grew very slowly over time. In contrast, the standard model's prediction drifted significantly, becoming unreliable after just a few steps. Specifically, the new method reduced the prediction error from a value of 0.361 down to 0.104. This improvement in accuracy translated directly into better performance. When the robots were asked to actually perform the tasks, the new system succeeded 78.2 percent of the time, a substantial jump from the 53.6 percent success rate of the standard visual-only model. The improvement held true across six different types of interactions, showing that the method works for a broad range of physical challenges.
One of the most important findings of the study is what happens when the robot stops using the physical data. The physical information is only available while the robot is in the training phase; it is not something a real-world robot can measure during a mission. The researchers found that once the robot finished learning with the help of the physical data, they could throw away the physical sensors entirely. The robot then operated using only its camera, just like the standard models. Despite losing access to the physical data, the robot retained the improved ability to predict the future. This suggests that the robot had successfully internalized the rules of physics into its visual understanding. It learned to see the world in a way that naturally respects how objects move and interact, rather than just memorizing visual tricks.
The study also explored whether simply teaching the robot to recognize where objects are located would be enough to solve the problem. They tried a different approach where the robot was forced to guess the position of an object directly from a single snapshot of the image. While this method made the robot very good at identifying where an object was at a specific moment, it did not help the robot predict how that object would move in the future. The robot could still guess the position accurately, but its predictions about the future path of the object were just as unreliable as the standard model. This distinction is vital: knowing where something is right now is not the same as understanding how it will change. The new method succeeded because it focused on the relationship between actions and future states, using the physical data to guide the learning of that relationship, rather than just teaching the robot to label the current scene.
The researchers also tested what would happen if the connection between the visual and physical data was broken. In one experiment, they paired the visual video of one task with the physical data of a completely different, unrelated task. In this scenario, the robot's predictions became chaotic, and its ability to control the robot collapsed to a success rate of just 16.6 percent. This result proves that the system relies on the correct pairing of visual and physical information. It is not enough to have access to physical data; the robot must learn the specific link between what it sees and how that scene physically evolves. The study confirms that the improvement comes from learning the correct dynamics of the world, not just from having more data.
In the end, the work demonstrates a path toward more reliable robotic planning. By using a temporary, privileged view of the physical world to train the robot's visual system, the researchers created a model that is far better at imagining the future. The robot learns to anticipate the consequences of its actions with greater precision, reducing the risk of failure when it tries to manipulate real objects. While the current experiments were conducted in a simulated environment, the approach offers a clear strategy for teaching machines to understand the physical world not just as a collection of images, but as a system of moving parts governed by consistent rules. The robot does not need to carry heavy physical sensors into the real world; it only needs to have learned the right way to see.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.