Multi-Resolution Tactile Imitation Learning for Contact-Rich Robotic Manipulation
This paper introduces MiTaS, a multi-resolution tactile imitation learning framework that fuses data from RGB cameras, a GelSight Mini, and a high-frequency Evetac sensor via a transformer-based architecture to achieve significantly higher success rates in contact-rich robotic manipulation tasks compared to vision-only or standard visual-tactile baselines.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are trying to put a key into a very tight, rusty lock. If you only use your eyes, you might see the keyhole, but you won't feel the tiny bumps or the moment the key starts to slip. You'd likely jam the key and give up. Now, imagine if you had a super-powerful sense of touch that could feel not just where the key is, but also the tiny, rapid vibrations happening a thousand times a second as you wiggle it into place.
This paper introduces a robot system called MiTaS (Multi-Resolution Tactile Sensing) that does exactly this. It teaches robots to "feel" their way through tricky tasks by combining three different types of "senses," much like how humans use both their eyes and their sensitive fingertips.
The Three "Senses" of the Robot
The robot doesn't just have one camera on its hand; it has a special trio of sensors working together:
- The Wide-Angle Eye (RGB Camera): This is like a standard security camera. It sees the big picture, like where the table is and where the object is sitting. It's good at seeing the scene but bad at seeing tiny details or fast movements.
- The High-Definition Fingertip (GelSight Sensor): Imagine a tiny, soft, squishy camera attached to the robot's finger. When the finger touches something, this camera takes a super-clear photo of the exact shape of the object it's pressing against. It's like having a fingerprint scanner that can see the texture of a key or a gear. However, it takes photos at a normal speed, so it might miss very fast events.
- The Super-Fast "Strobe" Eye (Evetac Sensor): This is the secret sauce. It's an event-based sensor that doesn't take normal photos. Instead, it acts like a high-speed strobe light that only flashes when something changes. If the key slips even a tiny bit or vibrates, this sensor screams, "Hey! Something moved!" It captures thousands of these tiny changes per second, things the other two sensors would completely miss.
How They Work Together: The "Conductor"
The problem with having three different sensors is that they speak different languages and move at different speeds. The camera is slow, the GelSight is medium, and the Evetac is lightning fast.
MiTaS acts like a conductor in an orchestra. It has a special brain (a neural network) that listens to all three instruments at once.
- It takes the "slow" visual info and the "medium" texture info.
- It mixes in the "fast" vibration alerts from the Evetac sensor.
- It fuses them all into a single, super-smart understanding of what is happening.
Once the robot understands the situation, it uses a "flow-matching" policy. Think of this as a GPS that doesn't just tell the robot where to go, but calculates the perfect, smooth path to get there, adjusting in real-time based on what the sensors are feeling.
The Big Test: Five Tricky Tasks
The researchers tested this system on five difficult jobs that require delicate touch, such as:
- Assembling gears: Fitting teeth together without them getting stuck.
- Wiping a board: Pressing just hard enough to clean a line without tilting the sponge.
- Screwing in a lightbulb: Aligning the threads perfectly.
- Inserting a key: The classic "tight lock" challenge.
- Connecting a bulb: Aligning small pins into slots.
The Results:
- Vision-Only Robots: When the robot only used its eyes, it failed almost everything (only 31% success). It couldn't see through the "occlusion" (when the hand blocks the view) or feel the tiny misalignments.
- Standard Touch Robots: Robots with just the "fingertip" camera (GelSight) did better (54%), but they still struggled with fast, tricky moments like the key getting stuck.
- MiTaS (The Winner): By combining all three senses, the robot achieved an 80% success rate. It could feel the vibrations of a slipping key and adjust instantly, something the other robots couldn't do.
The "Cheat Code" Training Method
Here is a clever trick the researchers used. They trained the robot using all three sensors, but then they told the robot: "Okay, now go do the job, but you can't use the super-fast Evetac sensor anymore."
Surprisingly, the robot still performed better than if it had never seen the Evetac sensor at all. It was as if the robot had learned a "secret language" from the fast sensor during training, and even without it, it could "hallucinate" or guess the right moves based on what it learned. This is called co-training, and it boosted performance by over 10% in some tasks.
The Bottom Line
This paper shows that for robots to handle delicate, contact-heavy tasks (like fixing a watch or assembling electronics), they need more than just eyes. They need a mix of high-resolution touch (to see the shape) and high-speed touch (to feel the vibration and slip). By teaching robots to listen to both, MiTaS allows them to solve complex problems that were previously impossible for machines to handle reliably.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.