← Latest papers
💻 computer science

Open Sign Bibles (OSB): An Open-Access Multilingual Sign Language Bibles Dataset

This paper introduces the Open Sign Bibles (OSB), the largest legally unencumbered multilingual dataset of its kind, comprising over 700 hours of openly licensed Bible videos in 20 sign languages aligned with parallel text in over 800 spoken languages to facilitate multimodal retrieval and contrastive learning research.

Original authors: Colin Leong, Joshua Nemecek, Kavitha Raju, Joel Mathew, Vijayan Asari, Elisabeth Leong

Published 2026-08-20
📖 5 min read🧠 Deep dive

Original authors: Colin Leong, Joshua Nemecek, Kavitha Raju, Joel Mathew, Vijayan Asari, Elisabeth Leong

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Language is the primary tool humans use to share thoughts, tell stories, and pass down history. For most of the world, this happens through sound and speech. But for millions of people who are Deaf, language is visual, formed by the movement of hands, the expression of the face, and the position of the body. These are sign languages, distinct and complex systems with their own grammar and culture. Yet, in the digital age, these languages have been left behind. While computers can easily search through written text or understand spoken words, they struggle to "read" sign language. The technology needed to find a specific sign within a long video, or to connect a sign to a written sentence, has been difficult to build because there simply has not been enough data. Most existing collections of sign language videos are either too small, locked behind paywalls, or restricted by copyright, leaving researchers without the raw material needed to teach machines how to understand the visual world of signers.

A team of researchers has now released a massive new resource designed to change this: the Open Sign Bibles dataset. This collection brings together over 700 hours of professionally translated Bible videos in 20 different sign languages. Unlike previous datasets that were scattered across the internet or required special permission to use, this entire library is freely available for anyone to download and study. The project connects these videos to the text of the Bible in more than 800 spoken languages, creating a bridge between the visual and the written. By matching a video of a signer in American Sign Language, for instance, with the same story written in English, Spanish, or Hindi, the researchers have created a powerful tool for training computers to understand the relationship between a sign and its meaning.

The dataset is built from two main sources. The first comes from the Digital Bible Library, which provided thousands of videos featuring signers from various backgrounds, often accompanied by graphics or text on the screen. The second source is a set of raw studio recordings from India, showing a single signer against a plain background. In total, the collection includes more than 8,000 individual video clips. To make this data useful for computers, the team did not just dump the videos online; they carefully organized them. They linked every video to specific verses from the Bible, allowing a computer to know exactly which story is being told. For a smaller portion of the videos, they went even further, manually marking the exact moments when specific signs, like "God" or "day," appear on the screen. This level of detail turns a simple video archive into a structured learning tool.

The researchers used this new dataset to test how well current technology could find specific signs within these long videos, a task known as "sign spotting." They tried two different approaches. The first method relied on measuring the physical distance between the signer's hands and body in the video, comparing them to known examples. This approach did not work well; the computer got confused by the continuous flow of movement and could not reliably pick out individual signs. The second method used a more advanced system that learned to understand the meaning behind the movements, similar to how a human recognizes a concept rather than just a shape. This approach showed promise. It successfully identified certain signs, such as "day" or "people," with a high degree of accuracy. However, the system still struggled with others. It sometimes confused signs that looked similar, like "man" and "god," even when they meant very different things. It also failed to recognize signs that were performed differently than the examples it had seen, such as a left-handed signer or a sign used in a specific cultural context that differed from the standard dictionary definition.

The results highlight both the potential and the current limits of this technology. The system worked best when the video it was analyzing looked very much like the examples it had been trained on. When the signer moved quickly, used a different hand, or signed a word with a unique meaning, the computer often missed the mark. The researchers found that while the technology can reliably spot a few common signs, it is not yet ready to understand the full complexity of a continuous conversation. The dataset itself, however, is a significant step forward. By providing a large, open, and legally clear collection of multilingual sign language videos, the researchers have given the scientific community a solid foundation to build upon. They have shown that with the right data, computers can begin to learn the visual language of the Deaf, but they also made it clear that much more work is needed to handle the nuances of human expression. The path forward involves gathering more diverse examples and refining the tools to better understand the subtle differences between signs that look alike but mean different things.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →