A HamNoSys-Guided Dataset and Baselines for Fine-Grained Isolated Handshape Recognition in Sign Language
This paper introduces a large-scale, HamNoSys-guided benchmark dataset of 144,000 RGB images from 15 participants for 160 handshape classes, establishing subject-dependent and leave-one-subject-out baselines using various deep learning models to advance fine-grained isolated handshape recognition and sign language technology.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are trying to teach a robot to understand human sign language. It's not just about watching someone wave their hands; it's about decoding a complex, silent vocabulary where the shape of a hand, the angle of a thumb, or the bend of a finger can change a word entirely. Think of it like the difference between a "thumbs up" and a "thumbs down"—to a human, it's obvious, but to a computer, those two gestures are just a collection of pixels that look suspiciously similar. This is the world of "fine-grained handshape recognition," a field where scientists try to build digital dictionaries that can read the subtle, tiny differences in how people form signs. The goal is to create tools that can translate sign language into text or speech, helping bridge the communication gap for millions of people who are deaf or hard of hearing. But to teach a computer this skill, you need a massive library of examples, and until now, that library has been missing a crucial, organized section.
Enter a new study that acts like a master librarian for sign language. The researchers built a giant, carefully organized photo album containing 144,000 pictures of hands, all taken from 15 different people. They didn't just grab random photos; they used a special "rulebook" called HamNoSys (Hamburg Notation System), which is like a universal alphabet for hand shapes used by linguists to describe signs without being tied to any specific spoken language. They asked their participants to pose their hands to match 160 specific shapes from this rulebook, capturing every twist, turn, and finger bend.
The team then tested four different types of "student" computers to see how well they could learn from this album. Two students looked at the actual photos (like a human looking at a picture), while the other two students looked only at a skeleton map of the hand's joints (like a stick-figure drawing). They ran two different exams: one where the students studied and were tested on the same people (a friendly, easy test), and another where they had to recognize hands from people they had never seen before (a much harder, real-world test).
Here is what they found: When the computers were allowed to study the same people they were being tested on, the photo-reading students did very well, with the most advanced one (a model called ViT-B/16) getting about 86% of the answers right. However, when the test switched to "unseen people," the scores dropped significantly, with the best models only getting around 45% correct. This suggests that while computers are getting good at recognizing hand shapes in general, they still struggle to generalize when a new person with a different hand size or style shows up. The study also highlighted that some hand shapes are so visually similar—differing only by a tiny bend in a finger—that even the best models get confused, mixing them up like twins in a crowd.
Ultimately, this paper doesn't claim to have solved the problem of sign language recognition. Instead, it provides a solid, reproducible foundation—a new, high-quality dataset and a set of baseline scores—that other researchers can use to build better tools. It shows us exactly where the current technology stands and points out that the next big challenge is teaching computers to recognize these subtle hand shapes no matter who is making them.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.