← Latest papers
💻 computer science

A Visuo-Tactile Data Collection System with Haptic Feedback for Coarse-to-Fine Imitation Learning

This paper presents a visuo-tactile data collection system featuring a direct-drive gripper for natural haptic feedback and real-time temporal annotation, designed to generate contact-rich, multimodal demonstrations that enable high-quality coarse-to-fine imitation learning.

Original authors: Yeseung Kim, Nayoung Oh, Jun Park, Teetat Thamronglak, Daehyung Park

Published 2026-05-12
📖 4 min read☕ Coffee break read

Original authors: Yeseung Kim, Nayoung Oh, Jun Park, Teetat Thamronglak, Daehyung Park

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). ✨ This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to teach a robot how to perform a delicate task, like picking up a fragile egg or screwing a tiny bolt into a tight spot. To do this, the robot needs to "watch" a human do it first. This is called Imitation Learning.

However, the old ways of teaching robots have two big problems, kind of like trying to learn to drive a car while wearing thick gloves and looking at a map instead of the road:

  1. The "Gloves" Problem (Lost Touch): Most teaching devices use handles that disconnect the human's fingers from the robot's gripper. It's like trying to feel the texture of a strawberry through a thick winter glove. You can't feel the subtle pressure needed to squeeze just right without crushing it.
  2. The "Blurry Movie" Problem (Lost Structure): When a human teaches a robot, they do a mix of fast, rough movements (like walking across the room) and slow, precise movements (like threading a needle). Old systems just record one long, continuous video. They don't tell the robot which parts were the "rough approach" and which parts were the "precise finish."

The New Solution: A "Super-Sense" Teaching Glove

The authors of this paper built a new device to fix these issues. Think of it as a high-tech, super-sensory glove that acts as a bridge between a human's natural instincts and a robot's need for data.

Here is how it works, using simple analogies:

  • Direct Connection (No More Gloves): Instead of a handle that disconnects the hand, this device has jaws that you squeeze directly with your own fingers. It's like the difference between using a remote control to open a door versus turning the doorknob yourself. You feel the exact resistance and pressure immediately. If the object is hard, you feel it; if it's soft, you feel it. This lets you demonstrate the "perfect squeeze" naturally.
  • Eyes and Skin (Vision + Touch): The device has a camera on top (like a head-mounted GoPro) to see what's happening. But it also has "skin" on the jaws. This skin is a special sensor that can "see" the shape of the object it's holding, even parts that the camera can't see because the fingers are blocking the view. It's like having X-ray vision for your fingertips, allowing the robot to understand the 3D shape of the object perfectly.
  • The "Pause Button" for Structure (The Annotation Button): This is the clever part. The device has a button on the handle. While you are doing the task, you can press and hold this button to tell the system, "Hey, I'm doing the precise, careful part right now!" When you let go, the system knows, "Okay, now we are just moving quickly to get there."
    • The Result: Instead of one long, confusing video, the system creates a "highlight reel" that clearly separates the Coarse (fast, rough approach) from the Fine (slow, careful interaction).

Why This Matters

The paper claims that by combining natural touch, detailed 3D shape sensing, and real-time labeling of "rough" vs. "precise" moments, they can create a special dataset.

Think of it like giving a student a recipe that not only lists the ingredients but also highlights exactly when to stir gently versus when to whisk violently. This allows robots to learn complex tasks much faster and more accurately, specifically for jobs that require a lot of physical contact and delicate force control.

In short: They built a tool that lets humans teach robots with their natural sense of touch and clearly mark the "important moments," so the robot learns not just what to do, but how to do it with the right amount of care.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →