← Latest papers
💻 computer science

Humanoid Hanoi: Investigating Shared Whole-Body Control for Skill-Based Box Rearrangement

This paper presents a skill-based framework for humanoid robots that utilizes a shared, task-agnostic whole-body controller to achieve robust, long-horizon box rearrangement, addressing distribution shifts through data aggregation and validating the approach via the new "Humanoid Hanoi" benchmark in both simulation and on the Digit V3 robot.

Original authors: Minku Kim, Kuan-Chia Chen, Aayam Shrestha, Li Fuxin, Stefan Lee, Alan Fern

Published 2026-02-24
📖 5 min read🧠 Deep dive

Original authors: Minku Kim, Kuan-Chia Chen, Aayam Shrestha, Li Fuxin, Stefan Lee, Alan Fern

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). ✨ This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are teaching a robot to play a very difficult game of Tower of Hanoi, but instead of moving wooden disks with its hands, it has to pick up heavy cardboard boxes, walk across a room, stack them on top of each other, and do it all without dropping anything or falling over.

This paper, titled "Humanoid Hanoi," is about building a robot brain that can handle this complex, multi-step task for a long time without getting confused or crashing.

Here is the breakdown of their solution, using some everyday analogies:

1. The Problem: The "Specialist" vs. The "Generalist"

Imagine you hire a team of specialists to move boxes:

  • Specialist A is an expert at picking up boxes.
  • Specialist B is an expert at walking while carrying a box.
  • Specialist C is an expert at putting a box down.

In many robot systems, when the robot switches from picking to walking, it fires Specialist A and turns on Specialist B. The problem? Every time you switch specialists, there's a "handshake" moment where the robot might stumble. If the robot has to do this 20 times in a row (a "long horizon"), those tiny stumbles add up, and the robot eventually falls over.

The Paper's Idea: Instead of hiring different specialists, they built one "General Manager" (called a Shared Whole-Body Controller or WBC).

  • This General Manager is always on duty.
  • The high-level "skills" (Pick, Walk, Place) are just like instructions given to the manager.
  • The manager doesn't change their personality or how they walk; they just adjust their focus based on the instruction. This keeps the robot's balance consistent, no matter what it's doing.

2. The Challenge: The "Training Gap"

Here is the catch: You can train this General Manager to be great at walking, and great at picking up boxes, but you can't train them on every single possible combination of those tasks beforehand.

Imagine training a driver only on empty highways. Then, you ask them to drive a car while carrying a wobbly stack of Jenga towers. Even if they are a great driver, the car might sway differently because of the load. The robot's "General Manager" might get confused when the instructions change slightly, leading to a fall.

3. The Solution: "Practice Makes Perfect" (Data Aggregation)

The authors realized that to fix this, they needed to let the General Manager practice the combination of tasks, not just the individual tasks.

They used a clever trick called Rollout-Based Data Aggregation:

  1. They let the robot try to do the whole Tower of Hanoi game in a simulation.
  2. Whenever the robot successfully did a move (even if it was a bit wobbly), they recorded exactly what the robot's body was doing.
  3. They fed this "real-world practice data" back into the General Manager's training.

The Analogy: It's like a music teacher who teaches you to play scales (individual skills). But to get you ready for a concert (the long task), they make you practice the entire song over and over, recording your mistakes, and then teaching you how to fix those specific mistakes. The robot learns how to handle the "wobbly stack" because it has actually practiced it.

4. The Test: "Humanoid Hanoi"

To prove this works, they created a new benchmark called Humanoid Hanoi.

  • The Setup: Three towers. Three boxes of different sizes.
  • The Rules: You can only move one box at a time. You can't put a big box on top of a small one.
  • The Goal: Move all boxes from one tower to another, stacking them perfectly.

This isn't just a quick task; it takes minutes of continuous walking, lifting, and balancing. It's designed to break robots that aren't robust.

5. The Results

They tested their "General Manager" approach against the old "Specialist" approach on a real robot named Digit V3.

  • The Old Way (Switching Specialists): The robot got tired, stumbled, and failed after a few moves.
  • The New Way (Shared Manager + Practice): The robot was much more stable. It could handle the "wobbly" moments better because its "brain" had seen those specific situations during the practice phase.
  • Real World Success: On the actual physical robot, they achieved a 40% success rate on the full, difficult task. While 40% sounds low, in the world of complex robotics, getting a robot to walk, pick up, and stack boxes for 5 minutes straight without falling is a huge victory.

The Big Takeaway

The paper teaches us that for robots to do long, complex jobs, we shouldn't just stitch together small, perfect skills. Instead, we should give the robot a single, consistent brain and let it practice the whole job repeatedly to learn how to handle the messy, real-world errors that happen when skills are chained together.

It's the difference between a relay race where the baton is dropped every time a runner changes, versus a marathon runner who carries the baton the whole way, learning to run smoothly despite the weight.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →