← Latest papers
💻 computer science

HRDexDB: A Paired Human-Robot Dataset for Cross-Embodiment Dexterous Grasping

The paper introduces HRDexDB, a comprehensive dataset featuring 2,100 synchronized human and robotic dexterous grasping trials across 100 diverse objects, providing high-fidelity spatiotemporal ground truth to serve as a foundational benchmark for cross-embodiment manipulation research.

Original authors: Jongbin Lim, Taeyun Ha, Mingi Choi, Jisoo Kim, Byungjun Kim, Subin Jeon, Hanbyul Joo

Published 2026-06-23
📖 4 min read☕ Coffee break read

Original authors: Jongbin Lim, Taeyun Ha, Mingi Choi, Jisoo Kim, Byungjun Kim, Subin Jeon, Hanbyul Joo

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to teach a robot how to pick up a delicate teacup. You could show it a video of a human doing it, but there's a problem: human hands and robot hands are built differently. A human hand has five flexible fingers that bend in unique ways, while a robot hand might have four stiff fingers or a different shape entirely. If the robot just copies the human's movements exactly, it might drop the cup or crush it because its "fingers" can't bend the same way.

This paper introduces HRDexDB, a massive new "training library" designed to solve this exact problem. Think of it as a bilingual dictionary for hands.

Here is how it works, broken down simply:

1. The "Side-by-Side" Recording Studio

Most previous datasets were like watching a movie of a human picking up a cup, or a separate movie of a robot doing the same thing, but never together. HRDexDB is different. It records both a human and a robot picking up the same object at the same time.

  • The Setup: Imagine a room surrounded by 23 high-definition cameras (like a security system on steroids) plus cameras worn on the robot's "head" and the human's "head."
  • The Action: A human picks up one of 100 different objects (from a coffee mug to a screwdriver). Immediately after, a robot teleoperated by a human (using a special glove suit) picks up that same object, trying to match the human's intent but using its own mechanical hands.
  • The Result: The system creates a perfect, synchronized 3D movie of both actions happening in the same space, capturing every finger movement, the object's rotation, and even the "squeeze" force the robot feels.

2. Why This is a Big Deal (The "Translator" Problem)

The paper argues that robots need more than just a video of a human; they need a translation guide.

  • The Analogy: Imagine a human trying to write a letter with a giant, clumsy glove. They can't write the same way they do with bare hands. They have to adapt their grip.
  • The Dataset's Role: HRDexDB provides the data needed to teach robots how to adapt. It shows the robot: "When the human touches the cup here with their thumb, you should touch it there with your specific finger to get the same result."

3. What They Tested (The "Exam")

The authors didn't just collect the data; they used it to test if robots could actually learn from it. They ran two main experiments:

  • Experiment A: The Contact Map Transfer
    They took the "touch map" of where a human's fingers touched an object and tried to translate it into a "touch map" for a robot.

    • The Result: When the robot used the translated map (learned from HRDexDB) instead of just copying the human's map directly, it was much more successful at picking up the object without dropping it. It learned to "speak robot" while keeping the human's "intent."
  • Experiment B: The "Find the Match" Game
    They asked the system: "Here is a human picking up a weird-shaped tool. Can you find a robot in our database that knows how to pick up that same tool?"

    • The Result: The system successfully found robot grasps that matched the human's style, even for objects the robot hadn't seen before during training.

4. The "Hard Mode" Challenge

The paper also used this dataset to test how good current computer vision software is at seeing hands and objects when they are tangled together.

  • The Challenge: When a hand grabs an object, it blocks the view. This is "occlusion."
  • The Finding: The authors tested top-tier AI programs on their dataset and found that these programs struggled much more with HRDexDB than with simpler datasets. This proves that HRDexDB is a "hard mode" benchmark that pushes AI to get better at seeing through the mess of fingers and objects.

Summary

HRDexDB is a massive, high-quality library of paired videos showing humans and robots doing the same dexterous tasks. It acts as a bridge, allowing robots to learn from human dexterity not by blindly copying, but by understanding how to translate human strategies into their own mechanical reality. The paper claims this is the first dataset of its kind to offer this level of detail (3D motion, tactile sensors, and multiple robot hands) for 100 different objects.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →