← Latest papers
💬 NLP

YoNER: A New Yorùbá Multi-domain Named Entity Recognition Dataset

This paper introduces YoNER, a new high-quality, multi-domain Named Entity Recognition dataset for the Yorùbá language comprising 5,000 sentences across five diverse domains, alongside the release of the OyoBERT language model and a comprehensive benchmark demonstrating the challenges of cross-domain generalization and the superiority of African-centric models for this task.

Original authors: Peace Busola Falola, Jesujoba O. Alabi, Solomon O. Akinola, Folashade T. Ogunajo, Emmanuel Oluwadunsin Alabi, David Ifeoluwa Adelani

Published 2026-04-08
📖 5 min read🧠 Deep dive

Original authors: Peace Busola Falola, Jesujoba O. Alabi, Solomon O. Akinola, Folashade T. Ogunajo, Emmanuel Oluwadunsin Alabi, David Ifeoluwa Adelani

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to teach a robot to read a storybook. You want the robot to be able to point out specific things like names of people, names of places, and names of organizations (like a school or a company). This task is called Named Entity Recognition (NER).

For languages like English, we have millions of storybooks, news articles, and websites to teach the robot. But for Yorùbá (a major language spoken by over 50 million people in Nigeria and neighboring countries), the robot has very little to learn from. Until now, the only "books" available for the robot were mostly news reports and Wikipedia articles.

This paper introduces a new, much bigger library for the robot called YoNER.

Here is the breakdown of what the researchers did, using some fun analogies:

1. The Problem: The Robot Only Knows "News"

Imagine you trained a robot to recognize animals, but you only showed it pictures of lions and tigers from a zoo. If you then showed it a picture of a lion in the wild, or a tiger in a cartoon, the robot might get confused.

Similarly, previous Yorùbá datasets were like that "zoo." They only had News (very formal) and Wikipedia (encyclopedic). They didn't have:

  • Bible stories (ancient names and places).
  • Blogs (casual slang and internet talk).
  • Movies (dramatic dialogue and shouting).
  • Radio shows (mix of formal and casual).

The robot was failing when it tried to understand these different "worlds."

2. The Solution: YoNER (The New Library)

The researchers built YoNER, a new dataset that acts like a massive, diverse library. They collected about 5,000 sentences from five different "neighborhoods":

  • The Bible: Full of ancient names and places.
  • Blogs: Full of slang, jokes, and casual talk.
  • Movies: Full of dramatic dialogue.
  • Radio: A mix of news and chat.
  • Wikipedia: The formal encyclopedia.

They hired three native Yorùbá speakers to manually read these sentences and highlight the names, places, and organizations. It's like having three expert librarians double-checking the robot's homework to make sure it's 100% correct.

3. The Experiments: Testing the Robot

Once they had this new library, they ran three main tests to see how well the robot could learn:

Test A: Can a "News" teacher teach "Movie" students?
They trained a robot using only the News data (the old way) and asked it to read Movies and Blogs.

  • Result: The robot struggled. It was like trying to teach a student how to write a rap song by only giving them a textbook on legal contracts. The robot got confused by the slang and the drama.
  • Good News: It did okay when moving from News to Wikipedia, because both are formal and serious.

Test B: Does a "Small Sample" help?
They gave the robot just a tiny taste (200 sentences) of the specific domain (like just 200 movie sentences) to see if that small help would fix the problem.

  • Result: Yes! Even a small amount of specific practice helped the robot understand that specific world much better.

Test C: The "Specialist" vs. The "Generalist"
They compared two types of robots:

  1. The Generalist (Multilingual): A robot that speaks 100+ languages but isn't an expert in any single one.
  2. The Specialist (Monolingual): A robot they built from scratch that only speaks Yorùbá (called OyoBERT).
  • Result: The Specialist (OyoBERT) was better at understanding the specific cultural nuances of Yorùbá, especially in movies and radio. However, the Generalist was still very good at moving between different topics because it had seen so much data from other languages.

4. The Big Takeaways

  • One size does not fit all: You can't just train a robot on news and expect it to understand movies or the Bible. You need data from all those different "neighborhoods."
  • Local experts win: A model built specifically for Yorùbá (OyoBERT) performed better than generic global models, proving that we need to build tools specifically for African languages, not just translate tools from English.
  • The "Gap" is real: Moving from formal text (News) to informal text (Blogs/Movies) is the hardest challenge. The robot gets lost in the slang and the drama.

Why Does This Matter?

Think of AI as a new student in a classroom. If we only give them textbooks, they will fail the pop quiz on the playground slang. By creating YoNER, these researchers are giving the AI student a complete education—from the church to the cinema, from the radio to the blog.

They are also releasing the OyoBERT model (the specialist robot) to the public. This means other developers can now build better apps for Yorùbá speakers, like voice assistants that understand local slang or search engines that find the right movie characters, not just news headlines.

In short: They built a diverse dictionary and a custom-trained robot to help AI finally understand the rich, varied, and beautiful world of the Yorùbá language.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →