← Latest papers
💻 computer science

The Embodiment Gap in Robot Foundation Models

This survey identifies and analyzes the "embodiment gap" in robot foundation models—the critical disconnect between reusable representations and the specific work required to adapt them for execution on different physical robots—by mapping existing methods, categorizing adaptation strategies, and proposing a comprehensive reporting framework to better evaluate cross-embodiment generalization.

Original authors: Yukiyasu Domae, Keisuke Shirai, Hanbit Oh, Ryoichi Nakajo, Tomohiro Motoda, Koshi Makihara, Masaki Murooka, Takuma Yagi, Yoshiaki Bando, Ryo Hanai

Published 2026-08-20
📖 5 min read🧠 Deep dive

Original authors: Yukiyasu Domae, Keisuke Shirai, Hanbit Oh, Ryoichi Nakajo, Tomohiro Motoda, Koshi Makihara, Masaki Murooka, Takuma Yagi, Yoshiaki Bando, Ryo Hanai

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine a world where robots could learn from the vast library of human knowledge, reading instructions and watching videos to understand how to pick up a cup, open a door, or sort laundry. This is the promise of a new generation of artificial intelligence called robot foundation models. These systems are trained on massive amounts of data, combining language, vision, and movement to create a general understanding of how the physical world works. The hope has been that if you feed a robot enough examples, it will become a universal helper, capable of walking into any kitchen or workshop and figuring out what to do. But there is a stubborn problem that has kept these smart systems from becoming truly useful in the real world. A robot is not just a brain; it is a body. And while the brain might understand the concept of "grasping," the specific way a robot's hand moves, the pressure it applies, and how it reacts when it touches something depends entirely on the unique shape and mechanics of that specific machine.

This disconnect between a smart, reusable brain and a specific, physical body is what researchers call the "embodiment gap." It is the difference between a plan that works in theory and the messy, difficult work required to make that plan happen on a real machine. A new survey from scientists at the National Institute of Advanced Industrial Science and Technology in Japan examines this gap in detail. They argue that while we have made incredible progress in teaching robots to understand language and vision, we have often overlooked the engineering work needed to connect those ideas to a robot's actual muscles and joints. The researchers found that simply making models larger or feeding them more data does not automatically solve the problem of getting a robot to move safely and effectively on a new body. Instead, significant work remains to translate a general idea into a specific, stable motion.

The researchers organized their findings by looking at how different groups of scientists are trying to bridge this gap. They identified three main approaches. The first approach focuses on sharing high-level ideas, like the meaning of a task or the visual cues of an object. For example, a robot might learn from a video that a cup needs to be lifted. This is useful, but the robot still needs to figure out exactly how to move its arm to reach the cup without knocking it over, a task that depends heavily on the length of its arm and the shape of its gripper. The second approach tries to standardize the data itself, creating common formats so that experiences from one robot can be used to train another. While this helps robots learn faster, the researchers found that even with standardized data, the physical reality of a new robot often requires adjustments to how commands are sent and how movements are executed. The third approach attempts to teach robots to understand the relationship between different bodies, learning how to translate a movement made by a human hand or a different robot into a motion that works for the current machine.

Despite these clever strategies, the survey reveals a consistent pattern: the closer a method gets to the actual physical movement, the more difficult it becomes to reuse it across different robots. When a robot is learning to grasp an object, the difference between success and failure often comes down to tiny details like friction, the exact angle of contact, or how the robot recovers if it slips. These are not problems that can be solved by simply adding more data to a computer model. They require careful calibration, safety checks, and often, human intervention to fix things when they go wrong. The authors point out that many research papers report high success rates for these systems but fail to mention the hidden work required to achieve them. A robot might succeed in a task 90 percent of the time, but if that success required researchers to constantly reset the robot, manually adjust its camera, or stop it from crashing every few minutes, the system is not yet ready for real-world use.

To address this, the researchers propose a new way of reporting results. Instead of just listing a final success percentage, they suggest that scientists should also report the amount of work needed to get there. This includes details like how many times the robot had to be reset, how much human help was needed to keep it safe, and how the system was adjusted to fit the new body. They introduce a simple checklist and a visual tool called an "adaptation curve" to show how performance improves as more work is put into adapting the robot. This transparency is crucial because it allows engineers and the public to see the true cost of deployment. It shifts the focus from just building smarter brains to building systems that can reliably and safely operate in the messy, unpredictable real world.

The survey concludes that while the vision of a universal robot is within reach, the path forward requires a change in how we think about these systems. We cannot rely on scaling up data and models alone to solve the physical challenges of robotics. The future of robot learning must include the engineering work of connecting a shared brain to a specific body, ensuring that the robot can not only understand a task but also execute it safely and recover from mistakes. By making the hidden work of adaptation visible, the researchers hope to guide the field toward creating robots that are not just smart, but truly ready to work alongside us.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →