← Latest papers
💻 computer science

iMaC: Translating Actions into Motion and Contact Images for Embodied World Models

This paper introduces iMac, a novel embodied world model paradigm that replaces low-dimensional structured action vectors with raw visual images as native action tokens to achieve superior prediction accuracy, task success, and generalization across diverse robotic embodiments.

Original authors: Zhenyu Wu, Xiuwei Xu, Yukun Zhou, Yifan Li, Qiuping Deng, Xiaofeng Wang, Zheng Zhu, Bingyao Yu, Ziwei Wang, Jiwen Lu, Haibin Yan

Published 2026-06-09
📖 5 min read🧠 Deep dive

Original authors: Zhenyu Wu, Xiuwei Xu, Yukun Zhou, Yifan Li, Qiuping Deng, Xiaofeng Wang, Zheng Zhu, Bingyao Yu, Ziwei Wang, Jiwen Lu, Haibin Yan

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). ✨ This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to teach a robot how to fold a shirt or pick up a cup. To do this safely and efficiently, you want to test the robot's brain (its "policy") thousands of times before letting it touch a real object. Usually, you'd have to build a physical simulation, which is hard, or just let the robot fail in the real world, which is slow and risky.

This paper introduces iMaC (Images of Motion and Contact), a new kind of "crystal ball" for robots. Instead of just guessing what the future looks like, iMaC acts like a highly accurate movie director that simulates exactly what will happen if a robot moves its arm in a specific way.

Here is how it works, broken down into simple concepts:

1. The Problem: The "Black Box" of Robot Actions

Most current robot simulators are like a magician who waves a wand and says, "I will move the robot." They take a robot's command (like "move arm 5cm forward") and turn it into a tiny, abstract code number. They then feed this number into a video generator and hope the robot's arm actually moves correctly in the video.

The problem is that in the real world, being off by just a few centimeters can mean the difference between grasping a cup and knocking it over. Abstract codes are too vague to capture these tiny, critical details. The simulator often has to "guess" (or hallucinate) where the robot's hand ends up, leading to errors that pile up over time.

2. The Solution: Drawing the Future Instead of Guessing It

iMaC changes the game by refusing to guess. Instead of using abstract codes, it draws the future for the robot. It translates the robot's commands into actual pictures that show exactly where the robot's body will be and how close it will get to objects.

It does this using two special "instruction manuals" (which the paper calls Motion Images and Contact Images):

  • Motion Images (The "Where"):
    Think of this as a blueprint. Before the video starts, iMaC uses the robot's official instruction manual (called a URDF) to calculate exactly where every joint will be. It then renders a video of the robot's arm moving through space, as if it were a 3D animation.

    • Analogy: Instead of telling a painter, "Draw a car moving," iMaC hands the painter a photo of the car already in the exact spot it needs to be. The video generator just has to fill in the background.
  • Contact Images (The "Touch"):
    This is the secret sauce. It's like a heat map showing the distance between the robot's hand and the objects in the room.

    • Analogy: Imagine a "force field" visualization. One map shows how close the robot's hand is to a table (Robot-to-Scene), and another shows how close the table is to the robot's hand (Scene-to-Robot). This tells the simulator, "Hey, the hand is this close to the cup; if it moves 1mm more, it will touch." This ensures the simulation knows exactly when a collision or grasp happens.

3. The "Self-Teaching" Loop

When you watch a long movie, the end of one scene becomes the start of the next. In robot simulations, if the model makes a tiny mistake in the first second, that mistake gets bigger and bigger by the end of the minute.

iMaC solves this by practicing what it preaches. During its training, it doesn't just look at perfect, real-world videos. It generates a chunk of video, takes the end of that generated video, and uses it as the start for the next chunk.

  • Analogy: It's like a student who practices by taking a test, grading their own answers, and then using those (sometimes imperfect) answers as the starting point for the next practice test. This teaches the model to handle its own mistakes so it doesn't get confused when running a long simulation in the real world.

4. The Results: A Better "Test Drive"

The researchers tested iMaC on eight difficult real-world robot tasks (like moving objects or manipulating tools). They used it to "test drive" different robot brains (policies) to see which ones were better.

  • The Finding: The scores iMaC gave to the robot brains matched the robots' actual performance in the real world almost perfectly.
  • Why it matters: If a robot brain performs well in the iMaC simulation, it will almost certainly perform well in the real world. This means engineers can safely and quickly pick the best robot "brain" without needing to run thousands of expensive, physical trials.

Summary

In short, iMaC is a robot simulator that stops guessing and starts showing. By turning robot commands into clear pictures of movement and distance, it creates a highly reliable "virtual reality" where we can test robot skills safely, knowing that if it works in the simulation, it will work in reality.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →