← Latest papers
💻 computer science

HoMMI: Learning Whole-Body Mobile Manipulation from Human Demonstrations

HoMMI is a framework that enables robots to learn complex whole-body mobile manipulation tasks directly from portable, robot-free human demonstrations by bridging the human-to-robot embodiment gap through a specialized cross-embodiment hand-eye policy design.

Original authors: Xiaomeng Xu, Jisang Park, Han Zhang, Eric Cousineau, Aditya Bhat, Jose Barreiros, Dian Wang, Shuran Song

Published 2026-03-04
📖 6 min read🧠 Deep dive

Original authors: Xiaomeng Xu, Jisang Park, Han Zhang, Eric Cousineau, Aditya Bhat, Jose Barreiros, Dian Wang, Shuran Song

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). ✨ This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you want to teach a robot to do chores around your house, like picking up laundry, delivering a package, or unfolding a tablecloth. The hardest part isn't just moving the arms; it's moving the whole body (walking, bending, turning the head) while keeping its eyes on the prize.

For a long time, teaching robots this way was like trying to teach a human by making them wear a heavy, awkward suit and move a joystick. It was slow, expensive, and hard to scale.

Enter HoMMI (Whole-Body Mobile Manipulation Interface). Think of HoMMI as a "Magic Costume" that lets regular humans teach robots by simply acting out the tasks themselves, using a smartphone-based system, without ever touching a robot.

Here is how it works, broken down into simple concepts:

1. The Problem: The "Uncanny Valley" of Movement

The researchers started with a system called UMI, where a person holds two phone-like grippers with cameras on them. This is great for teaching a robot how to move its hands. But, it's like teaching a driver how to park a car while only looking at the rearview mirror. You miss the big picture: the sidewalk, the other cars, and where you need to walk to get there.

To fix this, they added a camera to the person's head (like a GoPro). Suddenly, the human could see the whole room. But this created a new problem: The Embodiment Gap.

  • The Visual Gap: Humans are taller than robots, and our arms look different. If you just copy the video from a human's head to a robot's head, the robot gets confused because the world looks different from its shorter perspective.
  • The Kinematic Gap: Humans have flexible necks and can look up, down, and spin around easily. Robots often have stiff, limited necks. If you tell a robot to "copy my head movement exactly," it might try to twist its neck into a shape that breaks it.

2. The Solution: The "Translator" System

HoMMI acts as a brilliant translator between the human teacher and the robot student. It uses three clever tricks to bridge the gap:

  • Trick #1: The 3D "Ghost" View (Visual Translation)
    Instead of feeding the robot raw video (which looks different on a human vs. a robot), the system turns the human's video into a 3D point cloud. It strips away the human's body and arms, leaving only the "ghost" of the objects and the room. This way, the robot sees the shape of the world, not the person in it. It's like looking at a blueprint instead of a photograph; the blueprint works for any house, regardless of who built it.

  • Trick #2: The "Look-At" Point (Movement Translation)
    Instead of telling the robot, "Move your neck to these exact coordinates," the system tells the robot, "Look at this specific spot in the room."

    • Analogy: Imagine a human teacher pointing at a bird in a tree. The robot doesn't need to copy the human's exact arm angle. It just needs to know where the bird is. The robot then uses its own stiff neck to figure out the best way to look at that bird without breaking its neck. This is called a "relaxed" action.
  • Trick #3: The "Gripper-Centric" Map
    The system forces the robot to think from the perspective of its hands, not its head. Even though the human is walking around, the robot's brain stays focused on "Where are my hands relative to the object?" This keeps the robot stable and focused, like a tightrope walker who keeps their eyes on the rope, not the crowd.

3. The "Brain" and the "Body"

Once the human demonstrates a task (like folding a shirt), the system learns a Policy (the brain). This brain decides what to do next based on the 3D view and the "Look-At" points.

But the brain needs a body to move. That's where the Whole-Body Controller comes in.

  • Analogy: Think of the Policy as the Conductor of an orchestra, and the Robot as the Orchestra. The Conductor says, "Play the violin here, move the bass there." The Controller is the Musician who has to figure out exactly which fingers to press and how to move their body to make that sound happen without hitting their neighbor.
  • The Controller ensures the robot doesn't fall over, doesn't crash into walls, and moves smoothly, even if the "Conductor" (the human's demo) was a bit wobbly.

4. The Results: Can It Actually Do It?

The team tested this on three tricky tasks:

  1. Laundry: Picking up a shirt, walking to a bin, and dropping it in. (The robot had to look around to find the bin).
  2. Delivery: Carrying a box across a big room to a cart. (The robot had to navigate a long distance and turn to find the cart).
  3. Tablecloth: Unfolding a mat on a table. (This required two hands to work perfectly together while the robot moved its body).

The Verdict:

  • HoMMI (The New System): Succeeded about 80-90% of the time. It could walk, look around, and use two hands effectively.
  • Old Systems (Just hands or just head): Failed miserably.
    • Just hands: The robot walked into walls because it couldn't see the room.
    • Just head: The robot couldn't grab things because it couldn't see the fine details.
    • Trying to copy head movements exactly: The robot got stuck or moved erratically because it tried to mimic human flexibility it didn't have.

The Big Picture

HoMMI is a breakthrough because it allows us to teach robots by just doing the task ourselves, wearing a simple headset and holding two phones. It translates our human movements into a language the robot understands, ignoring the differences in height and body shape.

It's like having a universal translator that lets a human teach a robot to be a helpful assistant, simply by showing it how to do the job, without needing expensive, slow, or dangerous robot training sessions. This brings us one step closer to having robots that can actually help us around the house.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →