A Situational Speech Synthesizer for Yoruba: System Design, Phonological Rule Architecture, and Orthographic Extensions for Contour
This paper presents TTSYoruba, a rule-based concatenative diphone speech synthesizer for the Yoruba language that utilizes a hand-crafted phonological rule system to generate audio from tone-marked text, resolves complex nasal and tonal ambiguities, and introduces standardized orthographic markers for contour tones.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
The Music of Words: Teaching Computers to Sing in Yorùbá
Imagine language as a song. In many languages, the melody is hidden; you can change the pitch of your voice, and the meaning of a word might shift, but the written letters stay exactly the same. But in Yorùbá, a language spoken by millions in West Africa, the melody is written right into the notes. The language uses "tones"—high, low, and rising or falling pitches—to tell you exactly what a word means. If you write the wrong note, you might accidentally say "dog" instead of "lion." This makes Yorùbá a unique puzzle for computers. Usually, when we ask a computer to read text aloud (a technology called Text-to-Speech, or TTS), the machine has to guess the melody based on patterns it learned from listening to thousands of hours of human speech. But for languages like Yorùbá, where the "sheet music" is already written in the text, we don't need a computer to guess. We just need a computer that can read the notes perfectly.
This is where the story gets interesting. For a long time, computers struggled to speak Yorùbá naturally because they didn't have enough recorded human voices to learn from. They were like a musician trying to learn a complex song by ear without sheet music, often getting the rhythm wrong. But what if we built a robot musician that didn't need to guess? What if we gave it a library of every possible musical note combination, recorded by a human, and a strict set of rules to know exactly which note to play next? This is the challenge tackled by a team of researchers who wanted to build a voice for a massive online dictionary of Yorùbá names. They needed a system that could take a written name, figure out the exact pitch and rhythm, and speak it out loud instantly, without needing a human to record every single name in the world.
The Robot Choir and the Magic of Rules
The paper introduces TTSYoruba, a clever, rule-based robot voice designed specifically for Yorùbá. Think of this system not as a brain that learns by listening, but as a very organized librarian with a giant box of pre-recorded sound clips. The researchers recorded a human voice saying every possible combination of a consonant and a vowel (like "ba," "de," "go") in five different musical tones. They ended up with a library of 651 tiny sound clips.
The magic happens in the rules. When you type a name like Adébọ̀wálé, the computer doesn't just play the clips in order. It acts like a conductor. It looks at the notes written above the letters (the tone marks) and checks the note that came before it. In Yorùbá, the pitch of a note often changes depending on the note before it. If a low note is followed by a high note, the high note might actually "slide" up, sounding like a rising tone. The system has a specific rule for this: if it sees a high note coming after a low one, it swaps the standard "high" sound clip for a special "rising" sound clip. It does this for falling tones, too. By following these strict musical rules, the computer stitches the clips together to sound like a flowing sentence rather than a robot clicking its tongue.
Solving the "N" Mystery
One of the trickiest parts of the puzzle was the letter n. In Yorùbá, the letter n is a shape-shifter. Sometimes it's a consonant at the start of a word (like in nọ́). Sometimes it's a complete syllable on its own (like ń). And sometimes, it's not a sound at all, but a marker that tells you the vowel before it is "nasal" (sounding through your nose). To a computer, this looks like a confusing mess of identical letters.
The researchers built a "Nasal Disambiguation" system to solve this. It's like a detective that asks three questions:
- Does the n have a musical note on top of it? If yes, it's a standalone syllable.
- Is the n sitting between specific vowels like i or u? If yes, it's a nasal marker, and the computer changes the sound of the vowel before it.
- Is it just a normal consonant? If so, treat it as the start of the next sound.
By applying these rules, the system correctly pronounces names that would otherwise sound wrong, distinguishing between a nasalized vowel and a syllabic n.
The New Musical Notation: Carons and Circumflexes
Here is where the paper makes a bold, creative move. The researchers discovered a problem: some names have a "contour tone" (a pitch that goes up or down) on a single vowel, but the traditional way to write this in Yorùbá is to double the vowel (like Níkẹ̀ẹ́). This looks strange to many native speakers and feels clunky.
The team proposed a new way to write these sounds using symbols that already exist on computer keyboards but haven't been used for this purpose in standard Yorùbá writing: the caron (a little checkmark like ǎ) for rising tones and the circumflex (a little hat like â) for falling tones.
Think of it like this: instead of writing a long, double-note chord as two separate keys, you press one key with a special "hat" on it. The computer knows that ǎ means "play the rising sound clip," and â means "play the falling sound clip." This allows people to type names exactly as they are commonly spelled, without the weird double vowels, while still getting the perfect musical pitch. The researchers tested this by asking 50 people to listen to names written in the old "double vowel" way and the new "hat" way. The result? The listeners couldn't tell the difference. They rated both versions as equally natural and easy to understand. In fact, some listeners even preferred the new "hat" version because it looked more like the way they see the words in their heads.
The Verdict: A Good Start, But Not Perfect
The team tested their system with 50 volunteers who listened to 10 different names each. The results were promising but honest. The computer was incredibly good at being understood (scoring high on "intelligibility"), meaning everyone knew exactly what word was being spoken. However, the "naturalness" score was a bit lower. The listeners described the voice as sounding a bit slow, like a teacher reading a textbook to a class, rather than a friend chatting casually. This is because the system stitches together tiny clips, and there is a tiny, unnatural pause between every syllable.
The paper is very clear about what it doesn't do. It doesn't handle long sentences with complex rhythms or the way pitch drops over a whole paragraph. It is designed specifically for short names. It also relies on the user typing the correct tone marks; if you type without them, the robot won't know the melody.
Ultimately, this paper shows that for languages where the music is written down, we don't need massive AI brains to learn the song. We just need a smart librarian with a good set of rules. The system proves that a rule-based approach can handle the complex musicality of Yorùbá, and the new "hat" and "checkmark" symbols offer a practical, keyboard-friendly way to write these sounds without changing how people spell their names. It's a solid, working foundation that could help build even better voices in the future.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.