← Latest papers
💻 computer science

Humanoid-DART: Humanoid Loco-Manipulation using Diffusion-guided Augmentation through Relabeling and Tracking

Humanoid-DART is a self-supervised framework that combines diffusion-based trajectory generation with reinforcement learning to enable humanoid robots to learn diverse loco-manipulation skills from sparse demonstrations by automatically exploring the goal space with minimal expert supervision.

Original authors: Pranav Debbad, Kanish Thiagarajan, Victor Dhédin, Shafeef Omar, Majid Khadiv

Published 2026-06-26
📖 4 min read☕ Coffee break read

Original authors: Pranav Debbad, Kanish Thiagarajan, Victor Dhédin, Shafeef Omar, Majid Khadiv

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). ✨ This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine teaching a humanoid robot to do complex chores, like picking up a box and moving it to a specific spot, or kicking a ball. Usually, to teach a robot this, you need to record hundreds of hours of humans doing these tasks perfectly. But that's expensive, slow, and hard to do for every possible variation (what if the box is heavier? What if the target is further away?).

The paper introduces Humanoid-DART, a clever system that teaches a robot to master these tasks starting with just a tiny handful of examples (as few as four). It does this by acting like a creative coach who doesn't just repeat what they've seen, but imagines new ways to do things and then practices them until they work.

Here is how the system works, broken down into simple parts:

1. The Two-Team Strategy

Instead of trying to teach the robot everything in one giant brain, Humanoid-DART splits the job into two specialized teams that help each other:

  • The "Dreamer" (The Diffusion Model): This is the creative planner. Its job is to imagine a path for the robot to follow. It looks at the few examples it has and starts "dreaming up" new, slightly different paths to reach new goals. Think of it like an artist sketching new routes on a map. At first, these sketches might be physically impossible (like walking through a wall or sliding on the floor), but they are great at exploring where the robot could go.
  • The "Athlete" (The Reinforcement Learning Policy): This is the physical executor. Its job is to take the "Dreamer's" sketch and actually try to run it on the robot. It acts like a strict coach. If the Dreamer sketches a path where the robot's foot slides or it falls over, the Athlete says, "Nope, that won't work," and fixes the movement to make it physically possible.

2. The "Practice Loop" (The Curriculum)

The magic happens in how these two teams talk to each other over time. The system runs in cycles, like a training camp:

  1. Pick a Goal: The system picks a target (e.g., "Move the box to this new spot").
  2. Dream: The "Dreamer" creates a rough plan to get there.
  3. Test: The "Athlete" tries to execute the plan in a physics simulator.
  4. Filter & Relabel:
    • If the robot falls or fails, the plan is thrown out.
    • The Secret Sauce: If the robot almost makes it (a "near miss"), the system doesn't treat it as a failure. Instead, it says, "Okay, you didn't reach the intended spot, but you successfully reached this spot." It relabels the attempt as a success for the new spot it actually reached.
  5. Learn: The "Dreamer" learns from these successful (or near-successful) attempts to get better at dreaming up plans for those specific spots. The "Athlete" gets better at executing them.
  6. Repeat: The system picks new, slightly harder goals based on what it just learned, expanding its skills like a snowball rolling down a hill.

3. The "Dual-Brain" Architecture

To make the "Dreamer" really good at handling both walking and holding objects, the paper gives it a special two-part brain structure:

  • The Navigator: Focuses on the big picture (where the robot's feet are going).
  • The Localizer: Focuses on the details (how the hands and joints move relative to the object).
    By separating these, the robot learns to coordinate its whole body without getting confused, ensuring that when it walks toward a box, its hand is ready to grab it.

4. The Results

The researchers tested this on a real robot (a Unitree G1) and in simulation. They started with just four basic examples of a robot pushing, kicking, handing off, or picking up a box.

  • The Outcome: The system didn't just copy those four examples. It learned to do these tasks for goals that were 4 to 5 times further away than the original examples.
  • Coverage: For the "pick-and-place" task, the robot learned to cover 96% of the possible target areas, whereas other methods only covered a tiny fraction.

Summary

Humanoid-DART is like a student who starts with a few textbook examples but learns to solve thousands of new problems by:

  1. Imagining new solutions.
  2. Trying them out and seeing what actually works.
  3. Celebrating "near misses" as new learning opportunities.
  4. Slowly building a massive library of skills without needing a human to record every single variation.

The paper concludes that this approach allows robots to learn complex, whole-body tasks from very sparse data, making the process of teaching robots much faster and more efficient than before.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →