← Latest papers
💬 NLP

From Speech to Text Corpora: Evaluating ASR-Based Data Acquisition for Low-Resource Fongbe and Hausa

This paper evaluates the efficacy of using Automatic Speech Recognition (ASR) pipelines to generate text corpora for low-resource West African languages, demonstrating that fine-tuning MMS-300M on Fongbe significantly reduces word error rates while revealing that Hausa transcriptions from existing models approach acceptable quality, whereas Fongbe outputs currently require further post-processing.

Original authors: Mahounan Pericles Adjovi, Victor Olufemi, Roald Eiselen, Prasenjit Mitra

Published 2026-06-23
📖 5 min read🧠 Deep dive

Original authors: Mahounan Pericles Adjovi, Victor Olufemi, Roald Eiselen, Prasenjit Mitra

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to teach a computer to read and understand two specific African languages: Fongbe (spoken in Benin) and Hausa (spoken across West Africa). The problem is that there are very few books, websites, or written documents in these languages available for the computer to study. It's like trying to teach someone to speak a language by only giving them a few pages of a dictionary, when they need a whole library.

The researchers asked a simple question: Can we use the computer's ability to "hear" and "transcribe" speech to build that missing library? Since there are thousands of videos on YouTube in these languages (news, music, lessons), maybe the computer can listen to those videos and write down what is being said, creating a massive text database out of thin air.

Here is how they tried to do it, using some creative comparisons:

1. The Two Languages: A Tale of Two Tones

Think of Hausa as a standard road. It's a busy, well-traveled path where the computer has already seen some signs before. It's not perfect, but the computer has a decent map.

Fongbe, however, is like a mountain trail with invisible stepping stones. It is a tonal language, meaning the pitch of your voice changes the meaning of a word (like how a musical note changes a song). If you step on the wrong stone (get the tone wrong), you fall into a different meaning entirely. Most standard computer "ears" are deaf to these pitch changes; they hear the words but miss the music, often confusing Fongbe with French.

2. The Training: Tuning the Radio

To fix the computer's hearing, the researchers had to "tune the radio" specifically for these languages.

  • For Fongbe: They took a powerful, general-purpose listening model (called MMS-300M) and gave it a special training course using 12.3 hours of carefully recorded, clean speech. Think of this as a student studying with a strict, perfect teacher for a few weeks.

    • The Result: On the test (using the clean recordings), the computer became a star student. It made very few mistakes (only about 9.5% errors), a huge improvement over the old 44% error rate. It learned to keep the "musical notes" (tones) correct.
  • For Hausa: Since Hausa is better supported, they didn't need to build a new teacher. They just used an existing, well-trained model (a version of Whisper) that was already good at Hausa.

3. The Real-World Test: From Studio to Street

This is where the experiment got interesting. The researchers took their tuned models and pointed them at 424 real-world YouTube videos (about 45 hours of content).

  • The Setup: They didn't just pick quiet studio recordings. They picked real videos: news broadcasts, music videos, cultural shows, and outdoor interviews. This is like asking the student to take a test in a noisy cafeteria instead of a quiet library.

  • The Outcome for Hausa (The Road): The computer did a "good enough" job. About 60% of the transcriptions were usable. It was like a tourist who got lost a few times but still found the destination. The text wasn't perfect, but it was good enough to start building a library.

  • The Outcome for Fongbe (The Mountain Trail): The computer struggled significantly. Even though it was confident in its answers, the quality was low. Only 20% of the transcriptions were acceptable.

    • The Trap: The computer was "confidently wrong." It would write down a sentence that sounded right to its ears, but because it messed up the tones (the musical notes), the meaning was completely different. It was like a musician playing the right notes but in the wrong key, making the song sound like a different song entirely.

4. The Big Lessons

The paper draws a few clear conclusions from this experiment:

  • Small Data Can Go Far: You don't need a massive library to train a computer for a difficult language. Just 12 hours of high-quality, clean speech was enough to make the computer excellent at reading clean Fongbe.
  • The "Confidence" Trap: For tonal languages like Fongbe, you cannot trust the computer's confidence score. Just because the computer says, "I'm 90% sure I got this right," doesn't mean it did. It might be very sure about the wrong meaning.
  • Content Matters: The computer did much better on educational videos and news (where people speak clearly) than on music videos (where instruments and background noise drown out the speech).
  • The Gap: There is a big difference between a computer working in a quiet lab and working in the wild. For Hausa, the wild is manageable. For Fongbe, the wild is still too chaotic for the current technology to handle without human help to fix the mistakes.

Summary

The researchers successfully built a bridge from "spoken videos" to "written text" for these two languages. For Hausa, the bridge is sturdy enough to walk across. For Fongbe, the bridge is wobbly; the computer can cross it, but it needs a human guide to fix the steps before it's safe to use.

They are releasing all their tools, the list of videos they found, and the text they generated so others can try to improve the bridge.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →