GPU-Parallel Multi-Task Reinforcement Learning with Demonstration Guided Policy Optimization
This paper introduces MT-Libero, a GPU-parallel multi-task reinforcement learning benchmark for structured manipulation tasks, and proposes DGPO, an on-policy demonstration-guided method that combines importance-weighted PPO with adaptive behavior cloning to effectively train heterogeneous task suites under sparse rewards.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are trying to teach a robot arm to do chores. In the past, if you wanted the robot to learn how to pour coffee, you would train it for coffee. Then, if you wanted it to learn how to open a drawer, you had to wipe its memory and start a completely new training session just for the drawer. It was like hiring a different specialist for every single job.
This paper introduces a new way to train robots that is faster, smarter, and more versatile. It does two main things: it builds a massive, high-speed training gym, and it invents a new teaching method that uses human "demonstrations" as a guide without forcing the robot to just copy us blindly.
Here is a breakdown of their approach using simple analogies:
1. The Problem: The "One-Task-at-a-Time" Bottleneck
Think of current robot training like a single-lane road. You can only send one car (one robot learning one task) at a time. Even though we have powerful computers (GPUs) that could handle thousands of cars, we are stuck sending them one by one. This is slow and inefficient.
2. The Solution Part 1: MT-Libero (The "Super Gym")
The authors built MT-Libero, which is like turning that single-lane road into a massive, multi-lane superhighway.
- The Analogy: Imagine a giant gym where 40 different types of robots are training simultaneously in the same room. Instead of building 40 separate gyms, they built one giant, shared gym where every robot has its own station but they all share the same equipment, the same coach, and the same schedule.
- How it works: They took a standard set of robot tasks (like stacking blocks, opening drawers, or moving objects) and programmed the computer to run all of them at the exact same time on a single graphics card.
- The Result: They can train a robot on 40 different tasks at once, 1,600 times faster than before, using a tiny fraction of the computer memory. It's like training a "generalist" robot that knows a little bit about everything, rather than a "specialist" that only knows one thing.
3. The Solution Part 2: DGPO (The "Smart Mentor")
Training a robot on 40 tasks at once is hard because some tasks are easy (like picking up a light block) and some are hard (like opening a sticky drawer). If the robot gets good at the easy ones, it might stop trying to learn the hard ones. Also, sometimes the robot gets stuck and doesn't know what to do.
This is where DGPO comes in. Think of it as a Smart Mentor who watches a human expert perform the tasks.
- The Old Way: Some methods just say, "Copy the human exactly." If the human makes a mistake, the robot copies the mistake. Other methods say, "Ignore the human and figure it out yourself," which takes forever.
- The DGPO Way: The mentor says, "I'll watch the human to get you started, but I'll let you figure out the rest."
- When the robot is struggling: The mentor leans in close and says, "Hey, look at what the human did here. Do that." (This is called Adaptive Behavior Cloning).
- When the robot is doing well: The mentor steps back and says, "Great job! Now try to improve on your own." (This is the standard Reinforcement Learning part).
- The "Fairness" Rule: The mentor also keeps a scoreboard. If the robot is failing at the "hard" tasks but acing the "easy" ones, the mentor forces the robot to spend more time practicing the hard ones. This ensures the robot gets good at everything, not just the easy stuff.
4. The Results: A "VLA-like" Robot
The paper tested this system on a variety of tasks (Goal, Object, Spatial, and Long-horizon tasks).
- The Outcome: The robot trained with MT-Libero and DGPO became a true "generalist." It learned to do all 40 tasks in a single training session.
- Comparison: It beat robots trained without human help (which were slow) and robots trained by just copying humans (which were rigid and couldn't improve).
- Real-World Check: They even took a robot trained in this computer simulation and tried it on a real robot arm in a real room. It worked surprisingly well, showing that the skills learned in the "Super Gym" could transfer to the real world.
Summary
In short, this paper says:
- Stop training robots one by one. Build a shared, high-speed environment where they learn many tasks together (MT-Libero).
- Don't just copy humans; learn from them. Use human demonstrations as a flexible guide that helps when the robot is stuck but lets the robot improve on its own when it's ready (DGPO).
The result is a robot that learns faster, handles a wider variety of jobs, and gets better at the hard stuff without needing to be retrained from scratch for every new chore.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.