Wazobia Eval: A Benchmark for Nigerian Pidgin Emotion Understanding, Sarcasm Detection, and Cultural Reasoning
This paper introduces Wazobia Eval, a publicly available benchmark designed to address the underrepresentation of Nigerian Pidgin in AI evaluation by assessing emotion understanding, sarcasm detection, and cultural reasoning through a manually annotated dataset featuring a novel 16-category emotion taxonomy.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Language is more than a collection of words and grammar rules; it is a living map of a culture's history, its struggles, its humor, and its shared understanding of how the world works. For decades, the technology that powers our digital assistants and translation tools has been trained primarily on English and a few other major languages. These systems have become remarkably good at recognizing patterns in those languages, yet they often stumble when faced with the rich, context-dependent ways people speak in other parts of the world. This gap is particularly wide in West Africa, where Nigerian Pidgin serves as a vital bridge for tens of millions of people. While this language is spoken daily in markets, homes, and streets, the artificial intelligence systems designed to understand human communication have largely ignored its unique emotional landscape. The question facing researchers is not just whether a machine can translate a sentence from Pidgin to English, but whether it can truly understand what that sentence means to the person who said it.
To answer this, a researcher from Wazobia Labs in Lagos, Nigeria, introduced a new tool called Wazobia Eval. This is not a dataset of words to be memorized, but a test designed to measure how well a computer understands the emotional and cultural weight of Nigerian Pidgin. The researcher recognized that standard tests for artificial intelligence often rely on simple categories like "positive," "negative," or "neutral." However, in Nigerian Pidgin, a phrase like "I no fit shout again" (I cannot shout anymore) might not simply be negative; it could express a specific kind of exhaustion born from long economic struggle, a feeling that standard tests would miss. To capture this nuance, the researcher created a new framework containing sixteen distinct emotional categories. These include familiar feelings like joy and anger, but also culturally specific states such as "hustle fatigue," which describes the burnout from constant economic pressure, and "market energy," the sharp confidence needed for negotiation in a busy marketplace. They also included "forming," a term for when someone deliberately acts indifferent or composed to save face, and "prayer gratitude," a deep sense of thankfulness expressed through religious acknowledgment.
The researcher built this test using more than 550 examples of real Nigerian Pidgin conversations, carefully written and labeled by human experts who understand the local context. They did not just ask the computer to guess a feeling; they asked it to navigate complex social situations. One part of the test checks if the machine can detect sarcasm, where the words say one thing but the meaning is the opposite. Another part asks the machine to explain why a specific phrase was used in a certain context, testing its ability to reason about cultural norms rather than just matching keywords. For instance, the phrase "You don try well well" (You have tried very well) can be a genuine compliment, a sarcastic jab at a failure, or a sign of contempt depending entirely on who is speaking and what happened just before. The test requires the computer to know the difference.
When the researcher ran a preliminary test using GPT-5.5, the results were revealing. The model managed to correctly identify the emotion in less than half of the examples, achieving an accuracy of 43.75 percent. The errors were not random; they followed a clear pattern. The machine struggled most with the culturally specific categories and with phrases where the meaning depended entirely on the situation. It often confused sincere praise with sarcasm, or mistook a feeling of pride for simple happiness. These mistakes showed that the model was relying too heavily on the literal words rather than understanding the social context, the speaker's intent, or the shared experiences that give the words their true meaning. The study suggests that while modern language models are powerful, they still lack the deep cultural grounding required to truly understand how people in Nigeria communicate.
This work does not claim to have solved the problem of understanding African languages, nor does it suggest that current technology is useless. Instead, it provides a clear, reproducible way to measure where these systems fail and where they need to improve. By establishing a standard test that focuses on cultural reasoning and emotional nuance, the researcher has created a foundation for future development. Their goal is to ensure that as artificial intelligence becomes more integrated into daily life, from healthcare to education, it can serve African users with the same depth of understanding it offers to speakers of English. The paper concludes that progress in this field requires more than just more data; it requires a shift in how we evaluate machines, moving beyond simple translation to a genuine appreciation of the cultural realities that shape human language.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.