JSL-DC: A Word-Level Japanese Sign Language Dataset with Linguist-Derived Descriptions for Distinguishing Confusable Signs
This paper introduces JSL-DC, the largest Deaf-centric Japanese Sign Language dataset featuring 36.7K videos and linguist-derived descriptions for confusable signs, which enables a new model to outperform state-of-the-art recognition methods by 9.8% on challenging subsets.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
For many children who are deaf, the most critical years for language development happen before they ever enter a classroom. During this time, they rely on their parents to teach them how to communicate. Yet, statistics show that ninety-five percent of deaf children are born to hearing parents who do not know sign language. This creates a silent gap in the home, where parents struggle to connect with their children because they lack the vocabulary to express basic needs, emotions, or concepts. While online videos exist to help parents learn, true fluency often requires hands-on practice, a method that has proven difficult to scale without the right tools. To build effective learning aids, computer scientists need vast libraries of video data showing how different people perform sign language words. However, for Japanese Sign Language, such data has been scarce, often limited to a single person or a small set of words, making it impossible to train computers to recognize the language from a new user.
A team of researchers has now filled this void with the creation of JSL-DC, the largest collection of Japanese Sign Language videos ever assembled. The project was driven by a simple but profound realization: to teach a computer to understand sign language, the data must come from the people who use it every day. Instead of hiring actors or students to mimic signs, the researchers recruited nineteen deaf adults who use Japanese Sign Language as their primary mode of communication. These participants recorded themselves performing 270 specific words, chosen specifically to help hearing parents communicate with their young children. The resulting dataset contains over 36,000 videos, a volume three times larger than any previous Japanese sign language collection.
The true innovation of this work lies not just in the volume of data, but in how it was curated. The researchers worked closely with deaf linguists to select words that are most useful for early childhood communication, prioritizing verbs and adjectives over nouns, as these are the building blocks of early sentences. More importantly, the team identified pairs of signs that look nearly identical to the untrained eye but carry different meanings. In many cases, these signs differ only in subtle facial movements or mouth shapes, cues that standard computer vision models often miss. To solve this, the linguists provided written descriptions explaining exactly how to tell these confusing pairs apart. They noted, for instance, that the sign for "bitter" and the sign for "spicy" use the same hand movements but require different mouth shapes to distinguish them.
The researchers then tested whether these human descriptions could actually improve a computer's ability to recognize signs. They built a system that first identified the sign using standard video analysis, and then, if the sign was one of the confusing pairs, it examined the mouth movements specifically to make the final decision. This approach, inspired directly by the linguists' notes, significantly improved the computer's accuracy on the difficult signs. The system correctly identified the confusing words nearly ten percent more often than the best existing methods. This proves that incorporating human linguistic knowledge into machine learning models can bridge the gap where pure visual data fails.
The team also demonstrated how this technology can be applied in the real world by adapting a popular arcade game into a learning tool. In this version, players sign a word to release a colored ball into a field of other balls. If the computer correctly recognizes the sign, the ball matches the color and clears the screen; if the sign is wrong, the ball is the wrong color. This interactive game allows hearing parents to practice signing in a low-pressure environment, receiving immediate feedback on their performance. The entire dataset, along with the linguistic descriptions and the game software, is being released to the public to accelerate further research and development.
By centering the project on the deaf community, from the selection of the words to the final review of every video, the researchers ensured the data is authentic and reliable. Every video was checked twice by deaf experts to ensure the signs were performed correctly and that no personal information was accidentally captured in the background. This rigorous process filtered out nearly ten percent of the initial recordings, leaving a clean, high-quality resource. The work suggests that the future of sign language technology lies in collaboration between linguists and engineers, using human insight to guide machines. As these tools improve, they offer a tangible path toward closing the communication gap in homes across Japan, allowing parents to finally speak the language of their children.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.