← Latest papers
💬 NLP

Nsanku: Evaluating Zero-Shot Translation Performance of LLMs for Ghanaian Languages

This paper introduces Nsanku, a comprehensive benchmark evaluating the zero-shot translation performance of 19 large language models across 43 Ghanaian languages, revealing that while top models like Gemini-2.5-flash achieve moderate scores, no current model simultaneously demonstrates high performance and consistency, indicating they are not yet reliably usable for large-scale translation in these languages.

Original authors: Stephen E. Moore, Mich-Seth Owusu, Akwasi Asare, Lawrence Adu Gyamfi, Paul Azunre, Joel Budu, Jonathan Asiamah, Elias Dzobo, Kelvin Newman, Edmund O. Benefo, Gerhardt Datsomor, Onesimus Addo Appiah, A
Published 2026-05-07
📖 5 min read🧠 Deep dive

Original authors: Stephen E. Moore, Mich-Seth Owusu, Akwasi Asare, Lawrence Adu Gyamfi, Paul Azunre, Joel Budu, Jonathan Asiamah, Elias Dzobo, Kelvin Newman, Edmund O. Benefo, Gerhardt Datsomor, Onesimus Addo Appiah, Ama Branoa Banful, Lucas Woedem Kpatah, Saani Mustapha Deishini, John Ayernor

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

The Nsanku Report: Testing AI Translators on Ghana's Languages

Imagine you have a giant library of 19 different "super-brains" (AI models). Some are owned by massive tech giants, and others are open-source projects built by communities. You want to know: Can any of these brains translate English into the 43 different languages spoken in Ghana without having ever been taught those specific languages before?

This is exactly what the Nsanku paper did. The name "Nsanku" comes from the Akan language and means "musical instruments." Just as a band needs many different instruments to make music, this project needed many different AI models to test how well they handle the diverse "music" of Ghanaian languages.

Here is the story of what they found, explained simply.


1. The Setup: A Strict "Zero-Shot" Test

Think of these AI models as students taking a surprise exam.

  • The Rule: They were not allowed to study beforehand. They couldn't be "fine-tuned" (re-trained) on Ghanaian data. They had to rely entirely on what they already knew from their general training. This is called a zero-shot test.
  • The Test Material: The exam questions were 300 sentences from the Bible, translated into 43 different Ghanaian languages. The researchers used the Bible because it's one of the few places where you can find written versions of almost all these languages in one place.
  • The Grading: They used two different grading systems:
    • BLEU: Like a strict teacher checking if the student used the exact right words.
    • chrF: Like a more flexible teacher checking if the student got the general sound and structure of the sentence right, even if the exact words were slightly different.

2. The Results: Who Passed? Who Failed?

The "Star Students" (Proprietary Models)

Three big-name AI models from tech giants (Google, Anthropic, and OpenAI) came out on top.

  • Gemini-2.5-flash was the class valedictorian with the highest score.
  • Claude-sonnet-4-5 and GPT-4.1 were close behind.
  • The Analogy: These are like the students who went to the most expensive private schools. They have seen a lot of data and can guess the answers better than anyone else, but they still aren't perfect.

The "Community Students" (Open-Weight Models)

The rest of the models were open-source (free to use and modify).

  • The best of this group was kimi-k2-instruct, but it still scored significantly lower than the "Star Students."
  • The Gap: There is a clear gap between the expensive, private models and the free, community ones. The private models are currently much better at understanding these languages.

The "Language Difficulty" Factor

Not all languages were equally easy to translate.

  • Siwu was the "easiest" language for the AI to translate (highest score).
  • Nkonya was the "hardest" (lowest score).
  • The Twist: Surprisingly, the most widely spoken languages (like Twi) didn't always get the highest scores. Sometimes, languages with fewer speakers got higher scores. Why? Because the specific Bible translation used for those languages was clearer and more complete than the ones for the popular languages. It's like having a clearer map for a small village than for a big city.

3. The Big Problem: The "Unreliable Friend" Issue

This is the most critical finding of the paper. The researchers didn't just look at the average score; they looked at consistency.

  • The Analogy: Imagine you have a friend who is great at cooking Italian food but terrible at cooking Thai food. If you ask them to cook a random meal, you never know if you'll get a delicious dinner or a burnt mess.
  • The Finding: No single AI model was both "High Performing" AND "Consistent."
    • The best models were "High Performing but Inconsistent." They might translate Siwu perfectly but fail miserably on Nkonya.
    • The consistent models were "Consistent but Average." They gave the same mediocre result for every language, never failing badly but never doing well either.
    • The "Leaders" Quadrant: The researchers drew a chart with four corners. The top-right corner is the "Leaders" zone (High Quality + High Consistency). No model and no language ended up in this zone.

4. What This Means (According to the Paper)

The paper concludes that while these AI models are impressive, they are not yet reliable enough to be used for real-world tasks (like translating government documents, medical advice, or news) for Ghanaian languages.

  • The "Scriptural" Limit: The test was done using Bible verses. The authors warn that these models might do even worse on everyday conversation, news, or legal text because they haven't seen those types of words in their training.
  • The "Data" Problem: The low scores aren't because the languages are "hard" or "broken." It's because the AI hasn't seen enough examples of them. It's like trying to learn a language by reading only one book; you might get the gist, but you'll miss the nuances.

Summary

The Nsanku project built a giant scoreboard to test 19 AI models on 43 Ghanaian languages.

  1. Big Tech models are currently the best, but free models are catching up.
  2. Character-based grading (chrF) is a better way to judge these languages than word-for-word grading (BLEU).
  3. Most importantly: No AI is currently reliable enough to be trusted with these languages. They are like a student who sometimes gets an A+ and sometimes gets an F, depending on the specific language. Until we see a model that is consistently good, we cannot fully trust them for important tasks.

The paper has made all its data and code public so that researchers can keep testing and improving these models, hoping to eventually fill that "Leaders" quadrant.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →