Bridging Gaps in Natural Language Processing for Yorùbá: A Systematic Review of a Decade of Progress and Prospects
This systematic literature review analyzes a decade of research (2014–2024) on Yorùbá Natural Language Processing to identify critical challenges like data scarcity and tonal complexity, while highlighting emerging resources and techniques to guide future advancements for this and other under-resourced African languages.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine the world of Artificial Intelligence (AI) as a massive, bustling library. For a long time, this library has been stocked with millions of books in English, Chinese, and Spanish. The librarians (the AI models) have learned to read, write, and understand these languages perfectly because they have endless practice material.
But then, you walk over to the section for Yorùbá, a beautiful and complex language spoken by about 50 million people in Nigeria and beyond. You find that this section is almost empty. There are a few scattered pamphlets, some handwritten notes, and a few old textbooks, but no comprehensive library. This is the reality for many African languages in the world of Natural Language Processing (NLP)—the technology that helps computers understand human speech.
This paper is like a detective's report written by three researchers from the University of Limerick. They spent a decade (2014–2024) hunting down every single clue, study, and experiment ever done to teach computers how to speak and understand Yorùbá. They found 105 "cases" (studies) and pieced them together to tell us where we stand, what's working, and where the roadblocks are.
Here is the breakdown of their findings, using some everyday analogies:
1. The Language is a "Tonal Puzzle"
Imagine trying to learn a language where the meaning of a word changes entirely based on the pitch of your voice, like a musical instrument.
- The Challenge: In Yorùbá, the word ogun can mean "iron," "war," or "twenty," depending entirely on whether you sing it high, low, or mid-pitch.
- The AI Problem: Most computer programs are like people who are tone-deaf. They hear the letters but miss the music. If the computer doesn't get the tone right, it thinks you are talking about a war when you are actually talking about twenty dollars. This makes teaching the AI incredibly hard.
- The "Diacritic" Issue: It's like trying to read a recipe where someone forgot to write down the "s" or the "accent marks." In digital texts, people often skip the special marks (accents) on Yorùbá letters. The AI gets confused because "Ogun" (without the mark) looks different from "Ọgún" (with the mark).
2. The "Recipe Book" is Missing Ingredients
To teach a computer to cook (process language), you need a massive cookbook (a dataset) with thousands of recipes.
- The Shortage: For English, we have cookbooks with millions of recipes. For Yorùbá, we have maybe a few dozen.
- The Workarounds: Because the Yorùbá cookbook is so small, researchers have been trying a clever trick called "Transfer Learning." Imagine you are a master chef who knows how to cook Italian food perfectly. You try to learn how to cook Nigerian food. You don't start from scratch; you use your knowledge of heat, spices, and chopping (the general rules of cooking) and just learn the specific Nigerian ingredients.
- The Progress: Researchers have been using huge "global cookbooks" (multilingual models) and fine-tuning them for Yorùbá. It's not perfect yet, but it's better than starting with an empty kitchen.
3. What Have They Built So Far?
Despite the lack of ingredients, the researchers have managed to build some impressive tools:
- The Translators: They have built bridges to translate between English and Yorùbá. Some bridges are old and shaky (rule-based), while others are modern and smooth (neural networks).
- The Sentiment Detectors: They created tools that can read a tweet or a movie review and tell if the person is happy, sad, or angry.
- The Voice Actors: They have recorded thousands of hours of human voices to teach computers how to speak Yorùbá (Text-to-Speech) and how to listen to it (Speech-to-Text).
- The Dictionary Builders: They are working on tools that can identify names of people, places, and organizations in Yorùbá text.
4. The "Code-Switching" Hurdle
Imagine a conversation where two friends are speaking, but they keep swapping words between English and Yorùbá mid-sentence.
- The Reality: Younger generations in Nigeria often do this. They might say, "I went to the ọja (market) to buy suya."
- The AI Struggle: Computers get very confused by this mix. It's like a translator who expects a sentence to be in one language but gets a jumbled mix of two. This makes it hard for the AI to understand the true meaning.
5. The Cultural Shift
The paper points out a sad reality: Digital Desertification.
- Even though Yorùbá is a vibrant, living language, many young people are choosing to type and speak only in English online because the digital tools (keyboards, autocorrect, AI) don't support Yorùbá well.
- It's like a garden where the plants are dying not because they are weak, but because the gardener (the tech industry) hasn't built a proper irrigation system for them. If we don't build these tools, the language might disappear from the digital world, even if it survives in homes.
The Verdict: What's Next?
The researchers conclude that we have made great progress, but we are still in the "infancy" stage.
- We need more data: We need to record more voices and write more texts in Yorùbá.
- We need better tools: We need AI that understands tones and accents automatically.
- We need community: Just like a community garden, everyone needs to pitch in to grow these resources.
In a nutshell: This paper is a call to action. It says, "We have found the seeds (the research), and we have a few sprouts (the tools), but we need to build a greenhouse (better technology and resources) so that Yorùbá can thrive in the digital age, just like English and Chinese do."
The goal isn't just to make a computer speak Yorùbá; it's to ensure that 50 million people can use technology in their own language, preserving their culture and identity in the modern world.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.