← Latest papers
💻 computer science

Guided Discovery of New Behaviors using Diffusion Policies

This paper proposes a framework that combines Feynman-Kac correctors with a novel guiding potential and iterative trajectory optimization to systematically guide diffusion policies toward discovering diverse, executable behaviors that are often overlooked when training data is limited.

Original authors: Dian Yu, Sebastian Sanokowski, Majid Khadiv

Published 2026-06-09
📖 4 min read☕ Coffee break read

Original authors: Dian Yu, Sebastian Sanokowski, Majid Khadiv

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). ✨ This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have a robot that learned to do tasks by watching a human do them a few times. This robot uses a special "brain" called a Diffusion Policy. Think of this brain like a master chef who has memorized a few favorite recipes. If you ask the chef to cook dinner, they will almost always make their most famous dish (the "dominant behavior") because that's what they know best.

However, sometimes there are other valid ways to cook the same meal—maybe a different way to chop vegetables or a unique way to flip a pancake—that the chef never saw in the training videos. These are "rare behaviors." The problem is that the robot's brain tends to ignore these rare options, sticking only to the safe, common path.

The paper introduces a new method called GDNB (Guided Discovery of New Behaviors) to help the robot find these hidden, rare, but useful ways of doing things without breaking anything.

Here is how it works, step-by-step, using simple analogies:

1. The Problem: The Robot Gets Stuck in a Rut

When the robot tries to figure out what to do next, it usually picks the most obvious answer. It's like a GPS that only knows the main highway and refuses to show you the scenic backroads, even if the backroads are a valid way to get to your destination. The robot misses out on clever solutions because its training data was limited.

2. The Solution: A "Rare-Event" Treasure Hunt

The authors created a system to force the robot to look at the "backroads." They do this in three main stages:

Step A: The Compass (Guiding the Robot)

Usually, the robot follows a map based on what it has seen before. GDNB adds a special "compass" (called a Feynman–Kac corrector with a guiding potential).

  • The Analogy: Imagine the robot is walking through a foggy forest. Normally, it walks straight toward the trees it sees most often. The new compass doesn't tell it where to go, but it gently nudges the robot toward the "edge of the fog"—areas where the robot is less sure of itself.
  • The Goal: It encourages the robot to generate "frontier" ideas: actions that are a bit weird or rare, but not so crazy that they are impossible.

Step B: The Safety Net (Local Repair)

When the robot tries these rare, weird ideas, they often fail. The robot might try to grab a cup from the wrong angle and drop it.

  • The Analogy: Think of this as a safety net or a tuning fork. The robot proposes a wild idea (like "grab the cup by the handle while spinning"), and a specialized repair tool (called Sampling-Based Trajectory Optimization) quickly fixes the details. It tweaks the movement just enough so the robot doesn't drop the cup, turning a "failed weird idea" into a "successful weird idea."
  • The Result: The robot learns that the weird way of grabbing the cup actually works, as long as you adjust your grip slightly.

Step C: The Lesson Book (Retraining)

Once the robot successfully performs a rare, repaired action, the system saves this new success story and adds it to the robot's training library.

  • The Analogy: It's like the chef trying a new way to chop onions, realizing it works great, and then writing that new technique into their recipe book. Next time, the robot knows this new trick exists and can use it again.

3. The Results: Finding New Moves

The researchers tested this on robots doing tasks like pushing blocks, lifting boxes, and hanging tools.

  • What they found: The standard robot would always push a block from the same side. The GDNB robot discovered it could push the block from the other side, or flip a box before grabbing it, or release a tool in mid-air in a way no one had shown it before.
  • Real-world proof: They didn't just see this in a computer simulation; they tested it on real robots, and the robots successfully performed these new, rare moves in the real world.

Summary

In short, this paper teaches robots how to be creative rather than just memorizers.

  1. Nudge the robot to try rare, unusual ideas.
  2. Fix those ideas so they actually work.
  3. Teach the robot the new trick so it remembers it for next time.

This allows robots to discover multiple ways to solve a problem, making them more flexible and capable, even when they haven't seen every possible solution before.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →