Evaluating Large Language Models for Hausa and Fongbe Machine Translation: Benchmarks, Failures, and Metric Reliability
This study evaluates the translation capabilities of four large language models for English-to-Hausa and English-to-Fongbe, revealing significant performance disparities between the two languages, the unreliability of standard automatic metrics in correlating with human judgment, and the necessity of large-scale evaluation and multi-metric approaches for low-resource African languages.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are trying to translate a story from English into two very different African languages: Hausa (spoken by millions in Nigeria and Niger) and Fongbe (spoken by a smaller group in Benin). You have four powerful "robot translators" (Large Language Models like GPT-4, Claude, and Gemini) ready to do the job.
This paper is like a rigorous taste test to see which robot does the best job and, more importantly, whether the "scorecards" we use to grade them (automatic metrics) actually match what real human speakers think.
Here is the breakdown of their findings, using simple analogies:
1. The "Robot Race": Who Won?
The researchers asked the robots to translate thousands of sentences. The results were a tale of two very different worlds:
- Hausa (The "Good Enough" Runner): The robots did a decent job here. The translations were understandable and flowed well. It was like a runner finishing a race with a respectable time.
- The Winner: GPT-4o was the favorite among human judges, even though the computer scorecards gave the win to Claude.
- Fongbe (The "Stumbling" Runner): The robots struggled badly here. The translations were often nonsensical or broken. It was like a runner tripping over their own shoelaces.
- The Winner: Gemini was the least bad option, but even it only scored a "poor" grade from humans.
- The Big Gap: Hausa translations were roughly three times better than Fongbe translations. This wasn't because Fongbe is a "harder" language, but because the robots had seen way more training data for Hausa than for Fongbe.
Key Takeaway: Just because a robot is good at one African language doesn't mean it will be good at another. You can't assume a "one-size-fits-all" AI works everywhere.
2. The "Scorecard" Problem: Do the Computers Lie?
This is the most surprising part of the paper. The researchers compared the robots' performance against two things:
- Human Judges: Real native speakers rating the translations.
- Automatic Metrics: Computer algorithms (like BLEU, chrF++, BERTScore) that try to grade translations without a human looking at them.
The Mismatch:
- For Fongbe: The computer scorecards and the humans agreed. They both said, "Gemini is the best."
- For Hausa: The computer scorecards and the humans disagreed completely.
- The computers (specifically BLEU and chrF++) said: "Claude is the winner!"
- The humans said: "No, GPT-4o is the winner; Claude sounds robotic."
The "Broken Ruler" Analogy:
Imagine you are measuring the height of two people.
- For Fongbe, your ruler works perfectly.
- For Hausa, your ruler is bent. It tells you the shorter person is taller because it's measuring the wrong thing (like measuring how many letters are in a word instead of how well the sentence sounds).
3. The "Magic Mirror" That Reflects Nothing (Neural Metrics)
The paper tested a specific type of computer metric called BERTScore, which uses a "magic mirror" (an embedding model) to see if two sentences mean the same thing.
- The Glitch: For both Hausa and Fongbe, this mirror was broken. It looked at two completely different sentences and said, "These are 99% identical!"
- The Result: Because the mirror couldn't tell the difference between a good translation and a bad one, the scores were all the same (high numbers for everyone). It was like a judge giving every contestant a gold medal because they couldn't tell who was actually singing.
- The Lesson: You cannot trust these "smart" metrics for these languages until we fix the mirror.
4. The "Sample Size" Trap
The researchers tried testing the robots with small groups of sentences (500 or 1,000) and then with a huge group (10,000).
- The Trap: With small groups, the results were chaotic and misleading. For example, at 1,000 sentences, the data suggested Gemini was winning Hausa. But when they tested 10,000 sentences, it turned out Claude was actually winning.
- The Lesson: You need a large crowd (at least 2,500 sentences) to get a true picture. Small samples are like judging a whole movie based on a 10-second clip; you might get the wrong idea.
Summary of Recommendations
Based on this "taste test," the authors suggest:
- Don't trust the computer scorecards alone. Especially for Hausa, the computers got the winner wrong. You must have real humans check the work.
- Use the right ruler. If you must use a computer metric, use chrF++ (which counts character matches) rather than the "magic mirror" (BERTScore), which is currently broken for these languages.
- Pick your robot carefully. If you need to translate to Hausa, use GPT-4o or Claude. If you need to translate to Fongbe, use Gemini, but be prepared for poor quality.
- Don't use Fongbe translations for anything serious yet. The quality is too low; it would be like trying to build a house with wet sand.
- Hausa is ready for "data augmentation." The translations are good enough to be used as extra training data, provided you filter out the bad ones first.
In short: AI is getting better at translating African languages, but it's not there yet. The tools we use to measure its success are often broken, and we need real humans to keep the robots honest.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.