A HamNoSys-Guided Dataset and Baselines for Fine-Grained Isolated Handshape Recognition in Sign Language
This paper introduces a large-scale, HamNoSys-grounded dataset of 144,000 RGB images from 15 participants for 160 handshape classes and establishes baseline performance using various deep learning and machine learning models under both subject-dependent and leave-one-subject-out evaluation protocols to advance fine-grained sign language recognition.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Sign languages are full, complex languages with their own grammar and structure, used by millions of people around the world. Just as spoken languages are built from sounds, sign languages are built from physical movements of the hands, face, and body. One of the most critical building blocks is the shape the hand makes. A slight bend in a finger or a different position of the thumb can change the meaning of a sign entirely, much like how a single letter change can turn one word into another. For computers to understand or translate these languages, they must first learn to recognize these specific hand shapes with high precision. However, creating a computer system that can do this is difficult because human hands vary greatly from person to person, and the differences between similar hand shapes can be incredibly subtle.
Researchers have long sought a way to describe these hand shapes in a universal language that does not belong to just one country or culture. One such system, known as the Hamburg Notation System, acts like a detailed map for hand configurations, breaking them down into specific parts like finger selection and thumb position. While this system provides a clear definition for what a hand shape should look like, there has been a lack of real-world data to teach computers how to recognize these shapes when performed by actual people. Without a large collection of real images showing these specific shapes, computer programs struggle to generalize, often failing when they encounter a new person or a slightly different angle.
To address this gap, a team of researchers created a new, massive dataset designed to teach computers how to recognize 160 distinct hand shapes. They did not rely on computer-generated images or simple drawings; instead, they recorded real people. Fifteen volunteers, all university students, were asked to form each of the 160 specific hand shapes defined by the official chart. The researchers set up a controlled environment with a plain white background and a high-definition camera. Each participant held their hand in the required shape and slowly rotated it, allowing the camera to capture the movement from different angles. This process generated 144,000 individual color images, creating a rich library where every single picture is labeled with the exact hand shape it represents.
The researchers then tested how well different computer models could learn from this new library. They tried two main approaches. The first approach looked at the images as a whole, much like a human would, focusing on the overall appearance of the hand, skin tone, and lighting. The second approach ignored the picture entirely and focused only on the mathematical coordinates of twenty-one specific points on the hand, such as the tips of the fingers and the joints. They tested these models in two different ways. In the first test, the computer was shown images of people it had already seen during its training, simulating a scenario where the system knows the user. In the second, much harder test, the computer was shown images of people it had never met before, forcing it to rely on what it learned about the shapes themselves rather than the specific person making them.
The results revealed a clear distinction between knowing a person and knowing a shape. When the computer was tested on people it had already seen, it performed very well, correctly identifying the hand shape in more than 86 percent of cases using the best image-based model. However, when the computer faced a new person it had never encountered, its accuracy dropped significantly, falling to roughly 45 percent. This sharp decline shows that while computers are good at recognizing patterns in familiar faces, they still struggle to separate the universal shape of a sign from the unique way a specific individual performs it. The study also found that the models made the most mistakes with hand shapes that look nearly identical, differing only in a tiny detail like the bend of a single finger or the position of the thumb.
This work provides a solid foundation for future technology. By offering a large, balanced collection of real images and testing how computers handle new users, the researchers have created a standard tool for the field. The dataset allows scientists to build better systems that can eventually help translate sign language into text or speech for those who cannot hear, or help deaf individuals communicate more easily with hearing people. While the current models still find it difficult to recognize signs from strangers, the new resource gives researchers a clear path forward to improve these systems, ensuring that future technology can understand the subtle, beautiful complexity of human hands.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.