← Latest papers
💻 computer science

BEACON: Cross-Domain Co-Training of Generative Robot Policies via Best-Effort Adaptation

BEACON is a theory-driven framework that enhances the robustness and data efficiency of generative robot policies in cross-domain settings by casting co-training as a discrepancy-aware importance-reweighting problem, which jointly optimizes a diffusion-based visuomotor policy and per-sample source weights to minimize target-domain generalization error.

Original authors: Antong Zhang, Han Qi, Heng Yang

Published 2026-05-13
📖 4 min read☕ Coffee break read

Original authors: Antong Zhang, Han Qi, Heng Yang

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). ✨ This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to teach a robot arm how to stack blocks. You have two types of training data:

  1. The "Source" Data: A massive library of videos showing robots doing the task in a perfect, computer-generated world (simulation). There are thousands of these videos, but the lighting, textures, and physics are slightly different from the real world.
  2. The "Target" Data: A tiny, precious handful of videos showing a real robot doing the task in your actual kitchen. These are expensive and hard to get, but they are the only ones that truly matter for the final job.

The Problem:
If you just mix all the computer videos with the few real videos and teach the robot equally from all of them, the robot gets confused. The computer world is too different from the real world (different friction, lighting, camera angles). The robot might learn to stack blocks perfectly in the simulation but fail miserably in the real kitchen.

If you only use the few real videos, the robot doesn't learn enough and fails to generalize.

The Solution: BEACON
The paper introduces a new method called BEACON (Best-Effort Adaptation for Cross-Domain Co-Training). Think of BEACON not as a teacher, but as a smart editor or a curator.

Instead of treating every training video equally, BEACON assigns a "relevance score" (a weight) to every single video in the massive library.

  • The "Best-Effort" Idea: The system asks, "Which of these thousands of fake videos actually look and feel like the few real videos I have?"
  • The Editing Process:
    • If a computer video shows a robot moving in a way that is very similar to the real robot, BEACON gives it a high score. It says, "Use this one! It's helpful!"
    • If a computer video shows something weird or unrealistic (like a block sliding on ice when real blocks slide on wood), BEACON gives it a low score (or zero). It says, "Ignore this one. It will confuse the robot."

How It Works (The Magic Trick):
Usually, people try to force the computer world and the real world to look the same by changing the colors or textures (like putting a filter on a photo). BEACON does something different.

It doesn't try to make the worlds look alike. Instead, it selects the specific moments in the computer world that are already useful for the real world.

  • Analogy: Imagine you are learning to cook a specific dish. You have one recipe from your grandmother (the real data) and a thousand recipes from a famous TV chef who uses different ingredients (the source data).
    • Old Way: You try to force the TV chef's ingredients to match your grandmother's, or you just mix them all together.
    • BEACON Way: You read the TV chef's thousand recipes and pick out only the steps that match your grandmother's technique. You ignore the rest. You learn from the grandmother's few steps, but you fill in the gaps with the best matching steps from the TV chef.

The Surprising Result:
The paper found that by doing this "smart selection," the robot's brain (its neural network) naturally started to understand the real world better. Even though the system wasn't explicitly told to "align features" or "match patterns," the act of selecting the right data caused the robot to learn the right patterns automatically. It's like how a student who only studies the most relevant practice questions naturally starts to see the underlying logic of the subject, without needing a separate class on "logic."

What They Tested:
They tested this on a robot arm doing three tasks: stacking blocks, cleaning up mugs, and threading a needle. They compared BEACON against:

  1. Using only the few real videos.
  2. Mixing all videos equally.
  3. Using other methods that try to force the computer and real worlds to look similar.

The Outcome:
BEACON consistently won. The robot learned faster, used less real-world data, and was much more robust when the environment changed (like moving the camera or changing the table texture). It proved that being a smart curator of data is more powerful than trying to force the data to look the same.

In Short:
BEACON is a framework that teaches robots by saying, "We have a lot of fake data and a little bit of real data. Let's ignore the fake stuff that doesn't fit, and focus intensely on the fake stuff that does fit, so we can learn the real task efficiently."

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →