Design-Based Study of an Invisible Field Tool for Endangered Languages
This Design-Based Research study details the development and field testing of an Excel-VBA validation tool for the Mozilla Common Voice platform, demonstrating how iterative design improvements that prioritize community orthography preferences over technical features can enhance data quality and participant sustainability in endangered Circassian language projects.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
In the vast landscape of human communication, some languages are spoken by millions, while others hang by a thread, spoken by only a few thousand people scattered across the globe. When a language faces this kind of danger, the fight to save it often moves to the digital world. One of the most powerful tools for this purpose is a massive, global project where volunteers record themselves reading sentences aloud. These recordings are used to teach computers how to understand human speech, a technology known as speech recognition. For a language to be included in such a project, it needs a huge collection of these recordings, but gathering them is difficult when the speakers are few, live in many different countries, and may not be comfortable with complex computer software. The challenge is not just finding people to speak, but ensuring that the text they read is written correctly, a task that becomes incredibly hard when a language uses a unique alphabet that standard computer keyboards do not easily support.
This is the story of how a researcher tackled this problem for two endangered languages, Adyghe and Kabardian, which are spoken by Circassian communities in Russia, Turkey, and the Middle East. These languages share a deep history but were standardized into two separate written forms using the Cyrillic alphabet decades ago. Today, their speakers are dispersed, and many struggle with the specific characters required to write them correctly on a computer. A researcher named Mehmet Uğur Nemlioğlu set out to build a bridge between the community's desire to preserve their language and the technical demands of a global digital platform. He did not build a new website or a complex artificial intelligence system. Instead, he created a simple, invisible tool that lives inside a common spreadsheet program, designed to clean up text, fix errors, and organize the data so that ordinary volunteers could contribute without needing to be computer experts.
The journey began with a frustrating reality. Before this tool existed, the process of preparing sentences for the recording platform was slow and prone to mistakes. Volunteers would type sentences into shared documents, but these documents often contained hidden errors. The most common mistake involved a specific letter in the Cyrillic alphabet that looks like a vertical bar, used to mark a sharp stop in the throat. Because standard keyboards do not have this key, people often accidentally typed a regular Latin letter "I" or a number "1" instead. These tiny errors made the sentences unusable for the computers trying to learn the language. Furthermore, the process of checking these errors required someone with advanced technical skills to run complex command-line programs, a barrier that kept many community members from helping. The sentences would sit in limbo for days or weeks, waiting for a specialist to fix them, which often caused volunteers to lose interest and stop contributing.
To solve this, the researcher developed a tool that works inside a spreadsheet, a program that almost everyone already knows how to use. He built it piece by piece, testing it with the community and fixing problems as they appeared. The tool acts like a silent editor. When a volunteer pastes a list of sentences into the spreadsheet, the tool instantly scans every line. It finds the wrong characters and replaces them with the correct ones. It checks that punctuation is in the right place and that the sentences are not too long. It also sorts the sentences into different folders based on which version of the language they belong to, such as the version spoken in Turkey or the one spoken in Russia. This sorting is crucial because the two languages have different dialects, and the recordings need to be kept separate to be useful. Once the tool finishes its work, it produces clean files ready to be uploaded to the global platform in a matter of minutes, a task that previously took days.
The results of this approach were immediate and significant. The time it took to process a batch of one thousand sentences dropped from several days to roughly twenty to fifty minutes. This speed change did more than just save time; it changed the mood of the community. Volunteers could see their work move from preparation to the recording queue almost instantly, which kept them motivated to continue. The quality of the data also improved dramatically. The tool caught thousands of errors that would have otherwise slipped through, ensuring that the recordings uploaded to the global database were clean and accurate. The researcher observed that the tool became so integrated into the daily workflow that it felt invisible to the users, who simply saw their sentences getting fixed and organized without ever needing to understand the code behind it.
However, the study also uncovered a surprising human element that technology alone could not solve. The researcher had included a feature to automatically translate the Cyrillic text into a Latin alphabet version, hoping to help speakers who could not read the traditional script. He expected this to make the project more accessible. Instead, the community pushed back. Many speakers, even those who struggled with the Cyrillic script, felt a strong attachment to their traditional writing system. They preferred to read and record in the original script rather than a transliterated version. This resistance showed that in the effort to save a language, the community's feelings about their own identity and writing traditions are just as important as technical convenience. The researcher had to listen to this feedback and adjust the project, prioritizing the traditional script over the more "technically flexible" Latin option.
The success of this project offers a new way of thinking about how to build technology for endangered languages. It suggests that the most effective tools are not always the most complex ones, but rather those that disappear into the background, allowing people to focus on what they do best: speaking and preserving their culture. The researcher found that by removing the technical barriers, he could empower a small group of volunteers to do the work of a much larger team. Yet, this approach also came with a warning. Because the tool was so seamless, it was easy to miss subtle errors within the tool itself. It was only when the researcher stopped using the tool and looked at the code with fresh eyes, preparing to share it with the world, that he found small mistakes that had been hiding in plain sight for years. This taught him that even invisible tools need to be visible to the community so they can be checked and trusted.
Ultimately, this study demonstrates that saving a language in the digital age requires a partnership between technology and human trust. The tool did not just speed up a process; it rebuilt the relationship between the volunteers and the data. By making the work accessible, the researcher helped a scattered community feel connected and capable. The findings confirm that when technology is designed with the specific needs and feelings of a community in mind, it can become a powerful force for preservation. The path forward involves taking these lessons and building a web-based version of the tool that anyone can use, ensuring that the work of keeping these languages alive continues long after the initial project is finished. The story of Adyghe and Kabardian shows that the future of endangered languages lies not just in recording voices, but in creating the right space for those voices to be heard.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.