← Latest papers
💻 computer science

Human2Humanoid: Physics-Aware Cross-Morphology Motion Retargeting for Humanoid Robots

This paper introduces Human2Humanoid, an unsupervised framework that leverages a CycleGAN-based architecture with skeleton-aware graph convolutions and physics-aware constraints to achieve high-fidelity, physically plausible motion retargeting from humans to humanoid robots without requiring paired training data.

Original authors: Tianchen Huang, Feiyang Yuan, Junchi Gu, Shurui Fang, Xiaohu Zhang, Yu Wang, Wei Gao, Shiwu Zhang

Published 2026-06-03
📖 5 min read🧠 Deep dive

Original authors: Tianchen Huang, Feiyang Yuan, Junchi Gu, Shurui Fang, Xiaohu Zhang, Yu Wang, Wei Gao, Shiwu Zhang

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). ✨ This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have a talented human dancer and a brand-new robot that looks somewhat like a human but has different-sized arms, shorter legs, and joints that bend in different ways. You want the robot to copy the dancer's moves perfectly.

This is the problem the paper "Human2Humanoid" tries to solve. It's like trying to teach a toddler to dance exactly like a professional ballet dancer, even though the toddler has different proportions and can't bend their knees the same way.

Here is how the paper solves this, explained simply:

The Big Problem: "Apples and Oranges"

Usually, to teach a robot to move, you need a video of a human doing the move and a video of the robot doing the exact same move at the same time. This is called "paired data." But getting this data is incredibly hard, expensive, and often impossible because the robot might fall over if you try to make it copy a complex human move.

The authors say: "Let's skip the paired videos." They want to teach the robot using only videos of humans and only videos of robots, without ever seeing them move together.

The Solution: A "Translator" with Three Special Rules

The team built a smart AI system (a neural network) that acts like a translator between the human world and the robot world. To make sure the translation makes sense, they gave the AI three special rules:

1. The "Skeleton Map" Rule (Topology)

  • The Analogy: Imagine trying to draw a map of a city. If you just list street names, you might get lost. But if you know the shape of the city (where the rivers are, how the roads connect), you can navigate even if the street names change.
  • What the paper does: Humans and robots have different "skeletons" (different numbers of joints, different connections). The AI uses a Skeleton-Aware Graph Network. Instead of just looking at numbers, it understands the shape of the body. It knows that a human's elbow connects to a shoulder and a wrist, just like the robot's, even if the robot's arm is shorter. This helps the AI understand the structure of the movement, not just the raw numbers.

2. The "Relative Size" Rule (Morphology-Invariant)

  • The Analogy: If a giant and a midget both reach for a cookie on a high shelf, they both have to stretch their arms up. If you just told the midget to "reach 2 meters high," they would fail because they are small. But if you told them to "reach as high as your own head allows," they would succeed.
  • What the paper does: Humans and robots are different sizes. If the AI tried to match the exact coordinates (e.g., "move hand to X, Y, Z"), the robot would fail because its arms are shorter. Instead, the AI uses a Morphology-Invariant Loss. It looks at how far the hand moves relative to the body's starting pose (the "T-pose"). It tells the robot: "Move your hand the same percentage of your body length that the human moved theirs." This keeps the meaning of the move (like "reaching up") even if the robot is tiny.

3. The "Don't Fall Over" Rule (Physics-Aware)

  • The Analogy: If you tell a robot to walk, but you don't tell it to keep its feet on the ground, it might start "skating" on its toes or floating in the air like a ghost. It looks cool in a cartoon, but a real robot would crash.
  • What the paper does: The AI is forced to follow Physics-Aware Constraints. It has to check:
    • No Skating: If the human's foot is on the ground, the robot's foot must stay on the ground, not slide around.
    • No Floating: The robot's feet can't hover in the air when they should be touching the floor.
    • No Breaking: The robot's joints can't bend in ways that would break the machine (like bending a knee backward).
      The AI learns to "punish" itself if it generates a move that violates these rules.

The Result: A Robot That Can Dance

The team tested this on a real robot called the Unitree G1. They fed it human dance moves (without ever showing it a human-robot pair).

  • The Outcome: The robot successfully copied the human moves.
  • The Comparison: They compared their method to older ways of doing this (which rely on complex math equations to force the robot to fit). The new method was better at:
    • Tracking: The robot could actually follow the moves without falling.
    • Realism: The robot didn't slide its feet or float in the air.
    • Versatility: It worked on difficult moves like hopping, crouching, and turning, where older methods often failed or required manual tweaking.

In a Nutshell

The paper presents a new way to teach robots to move like humans without needing a "training partner" (paired data). It uses a smart translator that understands body shapes, respects size differences, and strictly enforces the laws of physics so the robot doesn't fall over or break itself. It's like teaching a robot to dance by showing it a video of a human, while the robot's brain automatically figures out how to adjust its own short legs and stiff joints to keep the rhythm without tripping.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →