Yorùbá in Unicode: An Overview of a Problem
This paper argues that the persistent failure of Yorùbá text rendering across digital platforms stems from Unicode's lack of precomposed characters for core diacritic combinations, and it proposes a formal encoding request to resolve these inconsistencies caused by the reliance on unstable combining sequences.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Language is more than just a collection of sounds; for millions of people, it is a system where the pitch of a voice changes the meaning of a word entirely. In the Yorùbá language, spoken by over 40 million people across West Africa and the diaspora, the difference between a high note and a low note can turn a word for "hand" into a word for "broom" or "respect." This reliance on pitch, known as tone, is not a poetic flourish but a grammatical necessity. Without the correct pitch markings, a simple sentence can become a confusing puzzle with dozens of possible meanings, leaving readers guessing whether a father moved inside a big farm or a big car. To solve this, writers use a system of diacritics—small marks placed above or below letters—to indicate these pitch changes.
However, the digital world was not built with these complex markings in mind. When computers first arrived, they were designed to handle standard letters from the English alphabet, not the layered symbols required for tonal languages. This created a persistent problem: while the technology could technically display the marks, it could not always keep them in the right place or count them correctly. A letter with a tone mark might look fine on one screen but turn into a blank box or a misplaced symbol on another. This paper by Kọ́lá Túbọ̀sún investigates why this happens, tracing the issue from the history of writing in Yorùbá to the rigid rules that govern how computers store text today. The author argues that the current system is fundamentally broken for tonal languages and proposes a specific, structural change to fix it.
The story of writing Yorùbá began long before the computer age. Originally an oral language, it was first written down by missionaries and traders in the 19th century using the Roman alphabet. Early writers struggled to adapt English letters to a language where pitch matters. Some tried to use extra letters, like adding an "h" or an "r" to show a tone, while others used dots or lines above vowels. Over time, a standard system emerged that used dots under certain vowels and marks above them to show pitch. By the mid-20th century, this system was well-established, allowing writers to distinguish between words that looked identical but sounded different. But when the digital age arrived, this carefully crafted system hit a wall.
The core of the problem lies in how computers store text. To make text travel easily between different devices and programs, a global organization called the Unicode Consortium created a standard list of characters. In this system, a letter with a single mark, like an "o" with an accent, is often stored as a single, pre-made unit. This works well for many languages. However, for Yorùbá, the system requires stacking two different marks on a single letter: a dot underneath to show the vowel type, and a mark above to show the pitch. The current digital standard does not have a single, pre-made code for these stacked combinations. Instead, it forces the computer to build them by sticking two separate pieces together: the base letter with the dot, and then the tone mark on top.
This method of building characters piece by piece is where the trouble starts. While it might work on one computer, the pieces can separate or shift when the text is moved to a different platform, a different font, or a different country. A writer might type a name perfectly, only to have the tone mark jump to the wrong letter or disappear entirely when the document is saved or printed. The author documents this failure across decades of publishing, showing book covers where the tone marks were hand-drawn after the text was printed because the typesetting machines couldn't handle them. In digital libraries, catalogers sometimes substitute the wrong symbols because the correct ones are missing, making it impossible for researchers to find books by their true titles. Even social media platforms, which count every character, treat these stacked marks as multiple separate characters, unfairly limiting how much Yorùbá text a person can write in a single post.
The paper challenges the official explanation for why these characters cannot be fixed. The Unicode Consortium maintains that their system is stable and that adding new pre-made combinations would break the consistency of text stored in older systems. They argue that the current method of combining pieces is sufficient and that the problems arise from poor fonts or old software. The author disputes this, pointing out that the failures happen even when the text is correctly encoded and moved between modern, high-quality systems. The issue is not just about how the text looks, but how it behaves: it fails to search correctly, it gets counted wrong, and it corrupts when transferred. The author notes that while the system has made room for thousands of new emojis in recent years, it has refused to update its rules for the essential needs of tonal African languages.
To solve this, the paper proposes a direct and specific intervention: the creation of four new, pre-made characters in the digital standard. These would be single units for the open "o" and "e" vowels with both a dot underneath and a pitch mark on top. This would allow the text to travel as a single, unbreakable unit, ensuring that a word for "hand" remains a word for "hand" no matter where it is read. The author acknowledges that current tools, like specialized keyboards and software that automatically add tone marks, have made writing easier, but they are temporary patches. They do not solve the underlying structural flaw that causes the text to break when it moves.
The conclusion is that the problem is not a lack of effort by writers or a failure of individual software, but a gap in the global rules that govern digital communication. The author argues that the Unicode Consortium has the technical ability to fix this and that doing so is necessary for the survival and proper use of the language in the modern world. By creating these four specific characters, the digital world could finally support Yorùbá in the same way it supports languages that do not require stacked marks. This would not only help the millions of Yorùbá speakers but would also set a precedent for other tonal languages facing the same digital barriers. The paper serves as both a record of the struggle and a clear roadmap for a solution that has been waiting for decades.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.