← Latest papers
💻 computer science

Bridging Handheld and Teleoperated Supervision for Contact-Rich Manipulation via State-Gated Experts

The paper proposes BRIDGE, a state-gated mixture of diffusion experts that effectively combines scalable handheld data with targeted teleoperated demonstrations to significantly improve success rates in contact-rich manipulation tasks by dynamically routing between action types based on the robot's current state.

Original authors: Vidullan Surendran, Neehar Peri, David Watkins

Published 2026-06-26
📖 4 min read☕ Coffee break read

Original authors: Vidullan Surendran, Neehar Peri, David Watkins

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). ✨ This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to teach a robot how to perform delicate tasks, like threading a needle or screwing a battery into a tight spring-loaded slot. You have two main ways to teach it, but both have a major flaw.

The Two Flawed Teachers

  1. The "Handheld" Teacher (The Fast Tourist):
    Imagine a human holding a controller that looks like a robot hand, walking around a room, and showing the robot what to do. This is fast and easy. You can collect hundreds of examples quickly.

    • The Problem: When the robot is just moving through empty air, this works great. But the moment the robot needs to touch something (like pushing a button or inserting a pipe), the human holding the controller can't feel the resistance perfectly. The human's hand might wiggle or push too hard because they are trying to "see" the path, not "feel" the physics. If the robot tries to copy this shaky movement exactly, it might smash the object or break the robot's arm.
  2. The "Teleoperation" Teacher (The Slow Surgeon):
    Imagine a human sitting at a high-tech console, controlling the robot directly with perfect precision, feeling every bit of resistance.

    • The Problem: This is perfect for the delicate parts, but it is incredibly slow and exhausting. It takes forever to collect enough data to teach the robot the whole task.

The Paper's Big Idea: "The Hybrid Coach"

The authors of this paper realized that you don't need a perfect teacher for every second of the task. You only need a perfect teacher for the tricky parts.

They came up with a system called BRIDGE (which stands for Bi-modal Routing for Imitation Data via Gated Experts). Think of it like a sports team with two specialized players and a smart coach:

  • Player A (The Handheld Expert): This player is great at running fast and moving through open space. They learn from the "Fast Tourist" data. They handle the boring, easy parts of the job.
  • Player B (The Teleop Expert): This player is a master of delicate, high-pressure situations. They learn from the "Slow Surgeon" data. They only step in when things get sticky.
  • The Coach (The Router): This is the magic part. The Coach watches the robot's current situation.
    • If the robot is just moving through the air? The Coach yells, "Player A, go!"
    • If the robot is about to touch a surface or push a spring? The Coach yells, "Switch to Player B immediately!"

Why "Naive Mixing" Fails

The paper tested what happens if you just throw all the data into a single pot and hope the robot learns the difference. It's like trying to teach a student by giving them a textbook on "how to run" and a textbook on "how to perform surgery" and telling them to just "figure it out." The robot gets confused. It tries to run with surgical precision (too slow) or perform surgery with running speed (too dangerous). The paper found that simply mixing the data actually made the robot worse than if it just used the handheld data alone.

The Results: Speed Meets Precision

By using their "Coach" to switch between the two players at the right moment, the robot achieved amazing results:

  • It learned the general shape of the task from the fast, cheap handheld data.
  • It only needed a tiny amount of slow, expensive "surgeon" data (just the tricky parts) to fix the mistakes.
  • The Outcome: The robot became up to 36.7% more successful at difficult tasks compared to using only the fast handheld data. It managed to get almost as good as the robot trained entirely by the slow surgeon, but it took a fraction of the time and effort to collect the data.

In a Nutshell

The paper solves a trade-off between speed and precision. Instead of choosing one or the other, they built a system that uses the "fast and dirty" data for the easy parts and the "slow and perfect" data only for the critical, contact-heavy moments. It's like hiring a tour guide for the walking tour but calling in a specialist mechanic only when the car breaks down.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →