← Latest papers
💻 computer science

SI-Diff: A Framework for Learning Search and High-Precision Insertion with a Force-Domain Diffusion Policy

This paper presents SI-Diff, a unified force-domain diffusion policy framework that integrates a novel mode-conditioning mechanism and a search teacher policy to simultaneously learn robust search and high-precision insertion, significantly improving misalignment tolerance and zero-shot transferability compared to existing baselines.

Original authors: Yibo Liu, Stanko Oparnica, Simon Shewchun-Jakaitis, Guoyi Fu, Jie Wang, Jun Yang, Anand Jagannathan, Tony Hong-Yau Lo

Published 2026-05-13
📖 5 min read🧠 Deep dive

Original authors: Yibo Liu, Stanko Oparnica, Simon Shewchun-Jakaitis, Guoyi Fu, Jie Wang, Jun Yang, Anand Jagannathan, Tony Hong-Yau Lo

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). ✨ This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to put a key into a lock, but you are wearing thick gloves and you can't see the keyhole very well. You have to feel your way around. If the key is slightly off-center, you might just push it straight down and get stuck. If you know exactly where the hole is, you can slide it in perfectly. But what if you need a robot to do this, and the robot doesn't know exactly where the hole is?

This paper introduces SI-Diff, a new "brain" for robots that solves this exact problem. It teaches a robot how to do two very different things using the same set of instructions: searching for the hole when it's lost, and sliding the object in once it's found.

Here is how it works, broken down with simple analogies:

1. The Problem: Two Different Moves

Usually, robots are taught to do these tasks separately.

  • The Search: If the robot is lost, it needs to wiggle and sweep the area (like a person feeling for a light switch in the dark).
  • The Insertion: Once found, it needs to be very delicate, wiggling just enough to slide the object in without getting stuck (like threading a needle).

Most robots have to switch "modes" or load different software to do these two things. The authors wanted to build a single robot brain that can do both without switching gears, just like a human worker who can feel around for a hole and then slide a peg in without thinking about changing tools.

2. The Solution: A "Diffusion" Policy

The team used a type of AI called a Diffusion Policy.

  • The Analogy: Think of diffusion like a sculptor starting with a block of noisy, random clay and slowly chipping away the noise until a perfect statue emerges.
  • In the Robot: The robot starts with a random guess of how to move its arm. The AI "denoises" this guess, refining it step-by-step based on what the robot's sensors (touch and force) are telling it, until it finds the perfect movement to either search or insert.

3. The Secret Sauce: The "Mode" Switch

The biggest challenge was teaching the AI to know which move to make (search vs. insert) without changing the code.

  • The Analogy: Imagine a music player that can play both a "Search Song" and an "Insertion Song." Usually, you'd need two different players. SI-Diff uses a Mode Prompt, which is like a "track switch" button.
  • How it works: The robot is told, "Right now, you are in Search Mode" or "Right now, you are in Insertion Mode." This tiny instruction tells the AI to look at the same sensor data but interpret it differently. In Search Mode, it looks for patterns that mean "keep looking." In Insertion Mode, it looks for patterns that mean "wiggle gently to fit."

4. The Teacher: Learning from a Smart Guide

Robots learn best by watching experts. The authors created a "Teacher Policy" (a set of rules) to generate training data.

  • The Search Teacher: Instead of just drawing a perfect spiral (which might miss the hole), this teacher is a bit chaotic. It tries eight different search patterns with random speeds and directions. It's like a teacher who doesn't just show one way to find a lost item, but shows you many different ways to sweep the floor so you learn to adapt.
  • The Insertion Teacher: This teacher knows how to wiggle a stuck peg out of a tight spot, a trick humans use instinctively.

The robot watches thousands of these "successful" attempts by the teacher and learns to copy the feeling of success.

5. The Results: From "Maybe" to "Definitely"

The team tested their robot against the current best methods (like a system called TacDiffusion).

  • The Old Way: If the peg was off by more than 2 millimeters (about the thickness of a credit card), the robot would usually fail. It was too rigid.
  • SI-Diff: The new system could handle misalignments up to 5 millimeters. It was much more forgiving of mistakes.
  • Zero-Shot Magic: Even though the robot was only trained on a square block, it could successfully insert unseen shapes like cylinders, triangles, and even a USB connector. It learned the concept of searching and inserting, not just the specific shape of the block.

Summary

SI-Diff is a robot control system that uses a single AI model to both hunt for a hole and slide an object into it. By using a "mode switch" and learning from a diverse set of "teacher" demonstrations, it makes robots much more robust, allowing them to handle bigger mistakes and work with shapes they've never seen before, all without needing to switch software or reprogram themselves.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →