← Latest papers
💻 computer science

HT-Bench: Benchmarking and Learning Dexterous Full-Hand Tactile Representations with Egocentric Vision

This paper introduces HT-Bench, a large-scale multi-task benchmark for dexterous full-hand tactile sensing, and proposes HandTouch, a vector-quantized vision-tactile encoder that significantly outperforms existing baselines in learning tactile representations through egocentric vision and full-hand tactile data.

Original authors: Yuzhe Huang, Jiaping Wu, Jiaming Jiang, Hezhe Lin, Aikebaier Aierken, Yunlong Wang, Kun Cheng, Ziyuan Jiao, Yuanxin Zhong

Published 2026-06-19
📖 5 min read🧠 Deep dive

Original authors: Yuzhe Huang, Jiaping Wu, Jiaming Jiang, Hezhe Lin, Aikebaier Aierken, Yunlong Wang, Kun Cheng, Ziyuan Jiao, Yuanxin Zhong

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). ✨ This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine trying to teach a robot hand to "feel" the world. For a long time, scientists have struggled to create a standard test for this, much like trying to compare the taste of apples from 200 different orchards when every orchard uses a different scale, a different type of fruit, and a different way of measuring sweetness.

This paper, HT-Bench, tackles that problem by changing the game. Instead of trying to force every robot hand to be the same, the authors created a massive, standardized "gym" specifically for robots that use eyes on the wrist (egocentric vision) and sensitive skin all over the hand (full-hand tactile sensing).

Here is the breakdown of their work using simple analogies:

1. The Problem: A Messy Kitchen of Sensors

Currently, robot tactile sensors are like a kitchen where everyone uses different measuring cups. Some sensors measure pressure, others measure slip, and they all come in different shapes. Because of this mess, it's hard to know if a robot is actually "learning" to feel or if it's just memorizing a specific trick for one specific sensor.

2. The Solution: HT-Bench (The "Gym")

The authors built HT-Bench, a giant dataset containing 10 million video frames and 7.8 million tactile frames. Think of this as a massive library of "movies" where a robot hand is doing 226 different tasks (like picking up a cup or turning a knob), and we can see both what the robot sees and exactly how its skin feels every single moment.

To test if a robot's "brain" is good at feeling, they don't just ask it to do one job. They give it four different types of exams:

  • The "Memory Match" Game (Retrieval): The robot is shown a picture of a hand touching something and asked, "Which of these 20 other hand-touches feels exactly like this one?" It's like a memory game where you have to match the feeling of a texture, not just the look.
  • The "Fill-in-the-Blanks" Puzzle (Inpainting): Imagine a hand touching a ball, but half the sensors on the hand suddenly go blind (like a patch of dead pixels on a screen). The robot has to guess what the missing sensors should be feeling based on the rest of the hand and what it sees with its eyes.
  • The "Mind Reader" Test (Vision-to-Tactile): The robot sees a picture of an object (like a sponge or a rock) but hasn't touched it yet. It has to predict exactly how the pressure will feel on its skin just by looking at the object. This is like seeing a picture of a lemon and instinctively knowing it will feel sour and bumpy before you even touch it.
  • The "Crystal Ball" Test (Prediction): The robot watches a hand moving and touching things over time. It has to predict what the hand will feel next second, based on what it saw and felt just a moment ago.

3. The New Brain: HandTouch

To pass these exams, the authors built a new AI model called HandTouch. They trained it in three progressive stages, like a student learning a new language:

  1. Stage 1 (Learning the Alphabet): The robot learns to look at a full map of its hand's touch and reconstruct it perfectly. It learns the "shape" of how pressure spreads across the hand.
  2. Stage 2 (Learning to Read Context): The robot is given a broken map (with missing parts) and a picture of the scene. It learns to use the picture to fill in the missing parts of the touch map. It learns that "if I see a soft pillow, my hand should feel soft, even if my sensors are broken."
  3. Stage 3 (Learning the Story): The robot learns to watch a sequence of events and predict the next feeling. It understands that if a hand is sliding down a ramp, the feeling will change in a specific way.

4. The Results: A New Champion

When they put HandTouch through the HT-Bench gym, it crushed the competition.

  • Memory Match: It got better at matching similar feelings than any previous model (improving from about 75% to 85% accuracy).
  • Fill-in-the-Blanks: It made far fewer mistakes when guessing missing touch data (reducing errors by more than half).
  • Mind Reader: It got much better at predicting how an object feels just by looking at it.

The Bottom Line

The paper concludes that while we can't yet build a single test for every robot hand in existence, we can build a massive, high-quality standard for robots that use eyes and full-hand sensors. HT-Bench provides the playground, and HandTouch proves that with the right training, robots can learn to "feel" the world in a way that is smart, adaptable, and generalizes to new tasks they've never seen before.

What they did NOT claim:

  • They did not say this robot can now perform surgery or handle delicate clinical tasks.
  • They did not claim this works on every type of robot sensor (like just a fingertip sensor or a force sensor on a leg).
  • They did not say this is ready for immediate commercial use in homes; it is a research benchmark to help future robots get there.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →