← Latest papers
💻 computer science

Hoi! - A Multimodal Dataset for Force-Grounded, Cross-View Articulated Manipulation

The paper introduces "Hoi!", a multimodal dataset comprising 3,048 sequences of articulated manipulation across 381 objects and 38 environments, which uniquely integrates visual, force, and tactile data from four distinct embodiments to bridge the gap between human and robotic interaction understanding.

Original authors: Tim Engelbracht, René Zurbrügg, Matteo Wohlrapp, Martin Büchner, Abhinav Valada, Marc Pollefeys, Hermann Blum, Zuria Bauer

Published 2026-04-17
📖 4 min read☕ Coffee break read

Original authors: Tim Engelbracht, René Zurbrügg, Matteo Wohlrapp, Martin Büchner, Abhinav Valada, Marc Pollefeys, Hermann Blum, Zuria Bauer

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). ✨ This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are teaching a robot how to open a stubborn, sticky drawer in your kitchen. If you just show the robot a video of you doing it, the robot might see your hand moving but have no idea how much push or pull is actually needed. It might pull too hard and break the handle, or not hard enough and fail to open it.

This paper introduces a new tool called Hoi! (short for "Hands-on Interaction") to solve exactly that problem. Think of Hoi! as a "Super-Teacher's Toolkit" for robots.

Here is the simple breakdown of what they did, using some everyday analogies:

1. The Problem: The "Silent" Video

Most robot training data is like a silent movie. You can see the actors moving (the video), but you can't hear the sound effects (the force).

  • Current Datasets: Like a cooking show where you see the chef chopping onions, but you don't know how much pressure they are using on the knife.
  • The Gap: Robots need to know not just what to do, but how hard to do it.

2. The Solution: The "Hoi! Gripper"

The researchers built a special robotic hand (a gripper) that acts like a super-sense human hand.

  • The "Magic Stick": Imagine a human holding a stick with a camera and a super-sensitive pressure sensor at the end. When they grab a drawer handle, the stick measures exactly how hard they are squeezing and pulling.
  • The "Digital Skin": The gripper has "fingertips" made of special gel that can "feel" the texture and pressure of the object, just like your skin does when you touch a door.

3. The Dataset: A "Multi-Camera, Multi-Sense" Movie Set

They didn't just record one person opening one drawer. They created a massive library of 3,048 video clips covering 381 different objects (fridges, dishwashers, cabinets) in 38 different rooms.

To make sure the robot learns everything, they filmed every single interaction from four different perspectives simultaneously:

  1. The Human Hand: Just a normal person opening a door.
  2. The "Wrist-Cam" Human: A person wearing a camera on their wrist, so the robot sees exactly what the human sees.
  3. The "Robot Hand" (UMI): A standard robot gripper doing the same task.
  4. The "Super-Hand" (Hoi! Gripper): The special sensor-equipped hand that records the force and touch data.

The Analogy: Imagine filming a magic trick. You have a wide-angle camera (exocentric), a camera on the magician's head (egocentric), and a camera on the magician's hand (wrist). But with Hoi!, you also have a "force camera" that records how hard the magician is pulling the rope. Now, the robot can learn the visual trick and the physical effort at the same time.

4. Why This Matters: The "Translation" Problem

The biggest goal of Hoi! is to help robots translate human skills into robot skills.

  • Before: If a robot sees a human open a fridge, it might try to copy the movement but fail because the fridge is heavy or stuck.
  • With Hoi!: The robot can look at the video, see the human's hand, and check the "force data" to say, "Ah, I see the human pulled with 15 Newtons of force. I need to do the same."

5. The Results: Robots Are Still Learning

The researchers tested existing AI models on this new dataset.

  • The Good News: The models got better at guessing what kind of object they were looking at (e.g., "That's a sliding drawer").
  • The Bad News: The models were still terrible at guessing the force. When they tried to predict how hard to pull, they were often wrong by a lot.
  • The Lesson: This proves that current AI is great at "seeing" but bad at "feeling." Hoi! provides the missing "feeling" data so future robots can learn to be gentle with fragile items and strong with heavy ones.

Summary

Hoi! is a giant, multi-sensory library that teaches robots the difference between watching a human open a door and feeling what it takes to open it. It bridges the gap between "seeing" and "doing," helping robots move from clumsy beginners to skilled, force-aware helpers in our homes.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →