Unified Motion Retargeting for Humanoids with Learned Point Cloud Correspondence
This paper presents Unified Motion Retargeting (UMR), a framework that learns dense point cloud correspondence to automatically and scalably transform diverse human motion data into high-fidelity robot trajectories without relying on manual human-robot mappings.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Humanoid robots are designed to move through our world with the same fluidity and adaptability that humans possess. To teach them how to walk, pick up objects, or navigate stairs, engineers often look to the vast library of human movement for inspiration. However, simply copying a human's motion onto a robot is not as straightforward as pressing a play button. Humans and robots are built differently; they have different body shapes, different numbers of joints, and different limits on how far those joints can bend. If a robot tries to mimic a human pose exactly, it might twist itself into an impossible position or lose its balance. For years, the solution has been to manually map specific human joints to specific robot joints, creating a custom instruction set for every new robot design. This process is slow, requires expert knowledge, and often fails to capture the subtle details of how a human body interacts with its environment.
A team of researchers has developed a new approach called Unified Motion Retargeting, or UMR, which changes how this translation happens. Instead of relying on a manual map of joints, the system treats the human body and the robot as two surfaces covered in millions of tiny points. Imagine the skin of a person and the skin of a robot as two clouds of dots. The researchers taught a computer to learn how these two clouds of dots correspond to each other, finding a match for every point on the human surface to a point on the robot surface. This learning happens once, in a neutral standing pose, and the resulting connection is saved. When the human moves, the system uses this pre-learned map to guide the robot's surface points to follow the human's surface points, rather than trying to match specific bones or joints. This allows the robot to reproduce complex poses and, crucially, to understand exactly where and how to touch objects or the ground, just as the human did.
The power of this method lies in its ability to handle differences in body shape without needing a new manual setup for each robot. In their tests, the researchers applied this system to a wide variety of human motion sources, including data from motion capture suits, scanned human meshes, and computer-generated animations. They then retargeted these movements onto five different humanoid robots, which varied significantly in height and body proportions. The system successfully transferred the motion across all these combinations, proving that the learned point-to-point map works regardless of the specific robot being used. This suggests that the method can scale to any new robot design without the need for engineers to spend weeks redefining how the robot's body parts relate to a human's.
Beyond just moving the body, the system also preserves the details of physical contact. When a human kicks a ball or climbs a staircase, specific parts of their body touch specific parts of the environment. Older methods often struggled to transfer these contact points accurately, leading to robots that might miss a step or fail to grasp an object correctly. Because the new system tracks the entire surface, it can transfer the contact map directly. If a human's hand touches a ball at a specific angle, the robot's hand is guided to touch the ball in the same way. In experiments involving tasks like carrying objects, kicking, and climbing stairs, the robots trained with this new method performed significantly better than those trained with previous techniques. They were more successful at completing the tasks and made fewer errors in their joint movements.
The researchers also tested the system in a real-world setting, deploying the learned behaviors on a physical robot. The robot was able to perform high-dynamic movements, such as a spinning kick, and navigate complex environments like stairs and slopes. The system handled these challenges with a level of stability and accuracy that matched or exceeded the best existing methods, even when the robot was operating in a simulated environment with random variations to test its robustness. The results indicate that by focusing on the surface geometry rather than the underlying skeleton, the system creates a more natural and reliable bridge between human motion and robot action. This approach offers a scalable path forward, allowing engineers to use large datasets of human movement to train a wide variety of robots without being limited by the specific design of the machine.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.