← Latest papers
🤖 AI

ExpertGen: Scalable Sim-to-Real Expert Policy Learning from Imperfect Behavior Priors

ExpertGen is a scalable sim-to-real framework that automates expert policy learning by initializing a frozen diffusion policy with imperfect behavior priors and refining it via reinforcement learning to achieve high success rates on complex manipulation tasks without reward engineering.

Original authors: Zifan Xu, Ran Gong, Maria Vittoria Minniti, Ahmet Salih Gundogdu, Eric Rosen, Kausik Sivakumar, Riedana Yan, Zixing Wang, Di Deng, Peter Stone, Xiaohan Zhang, Karl Schmeckpeper

Published 2026-03-18
📖 5 min read🧠 Deep dive

Original authors: Zifan Xu, Ran Gong, Maria Vittoria Minniti, Ahmet Salih Gundogdu, Eric Rosen, Kausik Sivakumar, Riedana Yan, Zixing Wang, Di Deng, Peter Stone, Xiaohan Zhang, Karl Schmeckpeper

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). ✨ This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you want to teach a robot how to do a complex task, like assembling a piece of furniture or stacking a banana on a can. Usually, you'd need to hire a human expert to show the robot exactly what to do thousands of times. But that's expensive, slow, and hard to scale.

EXPERTGEN is a new "robot teacher" system that solves this problem. It's like a smart, tireless coach that takes a few clumsy, imperfect attempts at a task and turns them into a master-level expert performance, all inside a virtual world before the robot ever touches the real world.

Here is how it works, broken down into three simple steps using a cooking analogy:

1. The "Bad Recipe" (Imperfect Behavior Priors)

Imagine you have a recipe for a cake, but it's a bit messy. Maybe it was written by a beginner, or maybe it was generated by an AI that hasn't baked much before.

  • The Problem: If you follow this recipe exactly, the cake might be lopsided, or it might fall apart. It's "imperfect."
  • The Paper's Solution: Instead of throwing the recipe away, EXPERTGEN treats this "bad recipe" as a starting point. It uses a special type of AI (called a Diffusion Policy) to learn the general "feel" of the recipe. It learns the basic motions, like "mix the batter" or "put it in the oven," even if the timing or temperature is slightly off.

2. The "Virtual Kitchen" (Massive Parallel Simulation)

Now, imagine you have a magical kitchen where you can run 1,000 simulations at the same time.

  • The Problem: If you just let a robot try to learn from scratch in this kitchen, it would crash into walls, drop the cake, and never figure out how to bake. It needs a guide.
  • The Paper's Solution: This is where the magic happens. EXPERTGEN takes that "bad recipe" (the imperfect prior) and puts it in the virtual kitchen. It uses a technique called Diffusion Steering.
    • The Analogy: Think of the robot's movements as a river flowing down a hill. The "bad recipe" sets the general direction of the river (it knows it needs to flow toward the ocean/task goal). However, the river might be too wide or hit some rocks.
    • The Fix: EXPERTGEN acts like a dam operator. It doesn't change the river's path entirely (which would make the robot forget the human-like style). Instead, it just nudges the water slightly at the very beginning of the flow to steer it around the rocks and toward the goal.
    • The Result: The robot learns to succeed by making tiny adjustments to its starting point, keeping the "human-like" style but fixing the mistakes. It does this millions of times in the virtual kitchen until it becomes a master chef.

3. The "Real World Transfer" (Sim-to-Real)

Now the robot is a master chef in the virtual kitchen. But can it cook in a real kitchen with real lighting, real dust, and real wobbly tables?

  • The Problem: Virtual worlds are perfect; real worlds are messy. A robot trained in a perfect simulation often fails when it sees a real object.
  • The Paper's Solution: EXPERTGEN uses a technique called DAgger (which sounds like a sword, but is actually a teaching method).
    • The Analogy: Imagine the virtual robot (the Master Chef) is teaching a student robot (the Student) who only has a camera and no access to the "secret state" of the kitchen.
    • The Process: The student tries to cook. When it gets confused or makes a mistake, the Master Chef steps in and says, "No, do it this way." The student records this correction. They repeat this over and over.
    • The Result: The student learns to handle the messy, real-world lighting and camera angles by constantly checking with the Master Chef. Eventually, the student becomes so good that it can cook perfectly on its own, even if the lighting changes or the table shakes.

Why is this a big deal?

  1. No "Reward Engineering" Headaches: Usually, teaching a robot requires a human to write complex rules like "If the robot drops the cup, give it -1 point." This is hard to get right. EXPERTGEN just says, "Did you finish the task? Yes? Good. No? Try again." It figures out the rest on its own.
  2. It Learns from Failure: Most robots give up when they fail. EXPERTGEN specifically learns how to recover from mistakes. If the robot knocks the banana over, it learns how to pick it back up and stack it, rather than just stopping.
  3. It's Cheap and Fast: You don't need to hire humans to record thousands of hours of video. You can generate the "imperfect" data with a computer script or a Large Language Model (LLM), and then let the simulation do the heavy lifting.

The Bottom Line

EXPERTGEN is like taking a rough draft of a robot's behavior, running it through a super-fast, infinite practice loop where it learns to fix its own mistakes, and then teaching a camera-based student how to copy that perfection in the real world.

The result? Robots that can handle messy, real-world tasks (like industrial assembly or picking up fruit) with a 90%+ success rate, even when they start with very little data and no human experts to guide them step-by-step.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →