A Single Diffusion-Policy Controller for Multi-Task Block Pushing with Zero-Shot Sim-to-Real Transfer
This paper proposes a framework that trains a single diffusion policy from scratch using reinforcement learning, enhanced by reverse curriculum generation and objective-centric representations, to successfully solve multi-task block-pushing problems in sparse-reward simulations and achieve zero-shot transfer to real-world environments with varying conditions.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you're teaching a robot arm to push a block across a table to a specific spot. Now, imagine that block isn't just a square; sometimes it's a "T," a "C," an "L," or even a weird "I" shape you've never seen before. The table might be sticky, slippery, or somewhere in between. The goal might be in a totally new place.
This paper is about teaching a robot to master all these messy, changing scenarios without needing a human to hold its hand and show it exactly what to do every single time.
The Big Idea: One Brain, Many Shapes
Usually, when robots learn to do tasks, they use a method called "Behavior Cloning." Think of this like a student copying a teacher's homework. The robot watches thousands of videos of experts pushing blocks and tries to mimic them. But here's the problem: if you want the robot to handle every possible shape and goal, you'd need a library of videos for every single combination. That's impossible to collect.
The authors argue that this "copying" method is too fragile. If you try to teach the robot a new task later, it often forgets how to do the old ones (a problem called "catastrophic forgetting").
Instead, this paper suggests a different approach: Reinforcement Learning (RL). Imagine the robot is a toddler learning to walk. It doesn't watch a video; it just tries, falls, gets a tiny "good job" (reward) when it gets closer to the goal, and tries again. The paper shows that by using a special kind of "brain" called a Diffusion Policy, the robot can learn from scratch, just by trial and error, to solve many different block-pushing puzzles at once.
The Secret Sauce: The "Reweighted" Brain
The core of their discovery is a new way to train this Diffusion Policy. In the world of AI, a "Diffusion Policy" is like a master artist who starts with a canvas full of static noise and slowly cleans it up until a perfect picture emerges. Usually, these artists are trained by looking at a gallery of perfect paintings (expert data).
The authors found a way to train this artist without the gallery. They created a simple rule: "If a move looks like it will get you a high score, make it more likely to happen." They did this by adding a special "weight" to the training process based on how good a move looks. It's like giving the robot a magic compass that points toward success, even when it's wandering in the dark.
They tested this against the old-school "Gaussian" policies (which are like robots that only guess the average, safe move). The results?
- The Old Way (Gaussian): In the simulation, the standard algorithms (like SAC, PPO, and TD3) mostly failed. They got stuck or gave up. One of them, SAC, managed to solve the task about 66.3% of the time, but it took a long time and was inconsistent.
- The New Way (ESM): The authors' new method, which they call ESM, solved the task 100% of the time in the simulation. Even better, it finished the job in an average of just 21.1 steps, whereas the next best method took nearly 119 steps.
The "Zero-Shot" Magic
Here is the most exciting part. The robot was trained entirely inside a computer simulation. It never saw a real block, a real table, or a real robot arm. Then, the authors took this digital brain and plugged it directly into a real robot in a lab.
They didn't retrain it. They didn't tweak it. They just turned it on. This is called Zero-Shot Sim-to-Real Transfer.
They tested the robot with:
- Different Friction: They put the block on a sticky floor mat and a slippery piece of parchment paper.
- Different Weights: They used a block that was only 1 cm thick (a quarter of the original 4 cm thickness).
- Different Shapes: They tested it on the shapes it learned in the simulator, plus a brand new "I-shaped" block it had never seen before.
The Results:
- On the real-world table, the new method (ESM) succeeded in 3 out of 3 trials for the T-block, C-block, and L-block.
- The old method (SAC) failed completely on the T-block (0 out of 3) and struggled with the others.
- Even on the brand new "I-shaped" block (which the robot never saw during training), the new method succeeded 61% of the time in the simulation, showing it could generalize to shapes it didn't know.
What Didn't Work (and What They Ruled Out)
The paper is very clear about what doesn't work well for this specific job.
- Just copying experts (Behavior Cloning): The authors argue that relying on massive datasets of expert demonstrations is a bottleneck. It's hard to get enough data for every possible shape and goal.
- Standard RL with simple guesses: They showed that standard algorithms using "Gaussian" policies (which assume the answer is usually just a simple average) just couldn't handle the complexity of pushing blocks with different shapes. They failed to learn the task in the simulation, let alone transfer it to the real world.
- Learning without a plan: When they removed their special "curriculum" (a training schedule that starts with easy goals and slowly gets harder) or their special way of describing the block's position, the robot's performance dropped significantly. This proves that the robot needs these specific training tricks to succeed.
How Sure Are They?
The authors are very confident in their simulation results, where they ran 4 million interactions to train the robot. They measured success rates and step counts with high precision.
When it comes to the real world, they tested the robot on 36 specific tasks across different conditions. They didn't just say "it worked"; they measured the distance the block traveled and the number of steps taken. They found that while the real world is messy, the robot's performance remained robust. However, they do admit that this "zero-shot" magic might not work for every type of robot task, especially ones that are much more complex than pushing a block on a flat table. For those harder tasks, they suggest that more work might be needed to bridge the gap between the computer and reality.
The Takeaway
This paper shows that you don't need a library of expert videos to teach a robot complex, multi-task skills. By using a smart, flexible "Diffusion Policy" trained with a new, simple reward system, a robot can learn to push blocks of all shapes and sizes, on slippery or sticky surfaces, and even handle goals it has never seen before—all without ever leaving the computer simulation until it's ready to work in the real world. It's a big step toward robots that can truly adapt to our messy, changing world.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.