← Latest papers
💻 computer science

One-Policy-Fits-All: Geometry-Aware Action Latents for Cross-Embodiment Manipulation

The paper proposes One-Policy-Fits-All (OPFA), a framework that learns a shared geometry-aware latent action space to enable a single policy to effectively manipulate diverse robot embodiments through end-to-end co-training, significantly improving data efficiency and cross-embodiment skill transfer.

Original authors: Juncheng Mu, Sizhe Yang, Hojin Bae, Feiyu Jia, Qingwei Ben, Boyi Li, Huazhe Xu, Jiangmiao Pang

Published 2026-03-17
📖 5 min read🧠 Deep dive

Original authors: Juncheng Mu, Sizhe Yang, Hojin Bae, Feiyu Jia, Qingwei Ben, Boyi Li, Huazhe Xu, Jiangmiao Pang

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). ✨ This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to teach a group of robots how to do chores. You have a robot with a simple two-finger clamp (like a clothespin), another with a three-finger gripper, and a third with a highly dexterous, human-like hand with five fingers and many joints.

The Problem:
Currently, if you want to teach the "clothespin" robot to pick up a banana, you have to collect hundreds of hours of video data just for that specific robot. If you then want to teach the "human-like hand" to do the same thing, you have to start from scratch and collect another huge pile of data. It's like hiring a chef who only knows how to cook with a wooden spoon, then hiring a new chef who only knows how to cook with a spatula, and forcing them to relearn everything from zero. They can't share their knowledge because their "hands" are built so differently.

The Solution: One-Policy-Fits-All (OPFA)
The paper introduces a new framework called OPFA (One-Policy-Fits-All). Think of OPFA as a universal translator and a shared brain for robots.

Here is how it works, using a simple analogy:

1. The "Shape-Shifting" Language (Geometry-Aware Latent Representation)

Instead of teaching the robots to move their specific joints (e.g., "move finger 1 up, move finger 2 down"), OPFA teaches them to think in shapes and spaces.

  • The Analogy: Imagine you are trying to describe a dance move to a friend.
    • Old Way: You say, "Move your left knee up 3 inches, then rotate your right ankle 15 degrees." This only works if your friend has legs exactly like yours.
    • OPFA Way: You say, "Reach out and grab the air in a circle." This instruction is about the shape of the movement, not the specific body parts.
  • How it works: The system looks at the robot's hand and converts its physical movements into a "point cloud" (a 3D map of where the hand can reach). It then compresses this 3D map into a single, abstract "idea" or latent code. This code represents the intent of the action (e.g., "grasp the object firmly") without caring if the robot has 2 fingers or 20.

2. The Universal Translator (Unified Decoder)

Once the robot has this abstract "idea" of what to do, it needs to translate it back into specific movements for its own body.

  • The Analogy: Think of the "idea" as a universal recipe (e.g., "Make a cake").
    • Old Way: You need a different recipe book for every type of oven (gas, electric, solar). If you want to use a new oven, you need a whole new book.
    • OPFA Way: You have one master translator. The robot takes the "universal recipe" (the latent code) and the translator instantly figures out, "Oh, your oven needs the dial set to 350, while that robot's oven needs it at 180."
  • The Magic: This translator is unified. It doesn't need to be retrained for every new robot. It can take the same "grasp" idea and instantly figure out how to move a simple gripper OR a complex human-like hand to achieve the same result.

3. The Super-Learning Effect (Cross-Embodiment Co-Training)

Because all robots are now speaking this "universal shape language," they can learn from each other instantly.

  • The Analogy: Imagine a classroom where a student with a bicycle, a student with a skateboard, and a student with a unicycle are all learning to navigate a maze.
    • Old Way: They practice separately. The skateboarder never learns from the bicyclist's experience.
    • OPFA Way: They all share a single "Maze Map" (the latent space). If the bicyclist learns a shortcut, the skateboarder and unicyclist instantly understand the concept of that shortcut, even though their vehicles are different. They just have to figure out how to apply it to their specific wheels.

Why This is a Big Deal

The researchers tested this on 11 different types of robot hands (from simple clamps to complex 5-finger hands) and found amazing results:

  1. Data Efficiency: Usually, teaching a new robot takes hundreds of examples. With OPFA, if you give a new robot just 8 examples (a tiny amount of data), it performs almost as well as a robot that has seen 72 examples. It's like learning a language by reading a few pages instead of a whole dictionary.
  2. Zero-Shot Transfer: A robot trained on data from a "two-finger" hand can successfully perform tasks in a workspace it has never seen, simply because it learned the geometry of the task from the other robots.
  3. No Custom Tuning: You don't need to write special code or retrain the system for every new robot you buy. You just plug it in, and the "Universal Translator" handles the rest.

In Summary:
OPFA stops robots from being stuck in their own "body silos." It teaches them to think about what they are doing (the shape and goal) rather than how their specific body parts move. This allows a robot with a simple gripper to learn from a robot with a fancy hand, making robot learning faster, cheaper, and much more adaptable to the real world.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →