Unadapted Multilingual ASR on a Garrusi Kurdish Evaluation Set: A Common-Reference Staged Normalization Analysis
This paper evaluates unadapted multilingual ASR models on a Garrusi Kurdish dataset by employing a common-reference staged normalization framework to demonstrate that orthographic mismatches between Latin-transliterated references and Arabic-script hypotheses significantly inflate error rates, while also revealing that pipeline limitations and model-specific shortcomings contribute to substantial residual errors.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Speech recognition technology has made remarkable strides, allowing machines to understand human voices across many languages. Yet, for speakers of less-documented dialects, this technology often hits a wall. The challenge is not just that a machine might fail to hear a specific sound correctly; sometimes the problem is that the machine and the human are speaking entirely different "languages" of writing. In the world of Kurdish, a language spoken across Iran, Iraq, Turkey, and Syria, this divide is sharp. Some varieties are written using a script based on Arabic letters, while others, or even the same variety in different regions, are written using a Latin alphabet with special marks to capture unique sounds. When a computer trained on one script tries to read a sentence written in another, it often fails not because it cannot hear the words, but because it cannot recognize the symbols. This creates a measurement problem: if a machine gets the sound right but writes it in the wrong script, standard tests count it as a total failure, hiding the fact that the machine actually understood the speech.
This is the specific puzzle Hiwa Asadpour tackled while studying Garrusi, a Kurdish variety spoken in Iran. Asadpour wanted to see how well a powerful, pre-existing speech recognition system could understand Garrusi speakers, even though the system had never been trained on that specific dialect. The system used, known as MMS-1B-all, was adapted for Central Kurdish and outputs text in the Arabic script. The researchers, however, had collected recordings from five Garrusi speakers and written down the correct words using a Latin alphabet with phonetic marks. To compare the two, they had to bridge the gap between the Arabic script the machine produced and the Latin script they had written down. They did this by creating a step-by-step process: first, they converted the machine's Arabic output into Latin letters, and then they simplified those letters to match the specific style of the reference text. This method allowed them to see exactly how much of the error came from the writing system difference and how much came from the machine actually failing to hear the words.
The results revealed a stark reality. When the researchers compared the machine's raw Arabic output directly against their Latin reference without any conversion, the system appeared to have failed completely. It scored a word error rate of over 111 percent, with zero words matching. This happened because the two scripts share no common letters; the machine was speaking a different visual language than the researchers. However, once the researchers converted the machine's output into Latin letters, the picture changed dramatically. The error rate dropped significantly, showing that the machine had actually recognized many of the sounds correctly, but the script mismatch had been masking this success. When they applied a final step to simplify the spelling differences, the error rate fell further, though it remained high. Even after fixing the writing system issues, the system still got only about 14 percent of the words exactly right. This means that while the writing system was a massive hurdle, the machine still struggled to understand the specific sounds and word structures of the Garrusi dialect.
The study also tested a different system that had been specifically fine-tuned for Southern Kurdish, a category that some linguists believe includes Garrusi. Surprisingly, this specialized system performed worse than the general Central Kurdish system. It produced far more extra words than the reference text and failed to match the speakers' actual words as well as the simpler system did. This suggests that simply training a model on a related dialect does not guarantee it will work well on a specific, distinct variety like Garrusi, especially when the data used for training comes from edited, read sentences rather than natural, field-recorded speech. The specialized system seemed to over-generate text, adding words that were not there, while the general system tended to miss words.
A key finding of this work is that the way we measure success matters just as much as the technology itself. The researchers showed that standard tests can be misleading if they do not account for differences in writing systems. By holding the reference text constant and only changing how the machine's output was presented, they could isolate the impact of the script conversion. They found that converting the script and simplifying the spelling reduced the measured errors by a large margin, proving that a significant portion of the "failure" was actually just a formatting issue. However, a large amount of error remained even after these fixes, indicating that the machine still has a long way to go to truly understand this specific dialect. The study concludes that for languages with multiple scripts or dialects, researchers must be careful about how they score results. They need to release the exact text they used for scoring so others can verify the numbers, and they must acknowledge that a high error rate might be partly due to the way the text is written, not just the machine's ability to hear.
This research highlights a broader issue in speech technology: the gap between high-tech models and the diverse, real-world languages they are meant to serve. For Garrusi speakers, and for many other communities with less-resourced languages, the path to accurate speech recognition is blocked not only by a lack of data but by the complexity of how those languages are written and classified. The study does not offer a perfect solution, but it provides a clearer map of the obstacles. It shows that before we can claim a machine understands a language, we must first ensure that the machine and the human are speaking the same visual language, and that our tests are fair enough to distinguish between a machine that cannot hear and one that simply writes differently.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.