SAFARI: A Community-Engaged Approach and Dataset of Stereotype Resources in the Sub-Saharan African Context
This paper introduces SAFARI, a community-engaged multilingual dataset comprising over 6,700 stereotypes across 16 languages from Ghana, Kenya, Nigeria, and South Africa, designed to address the underrepresentation of Sub-Saharan African contexts in generative AI safety assessments through culturally sensitive, native-language survey methods.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are teaching a new student how to understand the world. You give them a giant textbook filled with stories about people. But here's the catch: the textbook is mostly written by people from New York, London, and Mumbai. It has thousands of pages about life in those cities, but only a few blurry, outdated pages about life in Ghana, Kenya, Nigeria, and South Africa.
When this student (an Artificial Intelligence) tries to talk about someone from those African countries, it often makes things up or repeats harmful, old-fashioned rumors because it doesn't have the real, local knowledge.
This paper, titled SAFARI, is about fixing that textbook.
The Problem: The "Blind Spot"
Think of current AI models like a tourist who has only visited a few big cities. If you ask them about a village in the African savanna, they might guess based on movies or old stories. They might say, "Oh, everyone there is lazy," or "They are all rich," or "They practice magic." These are stereotypes—oversimplified, often unfair ideas about a group of people.
While researchers have been collecting lists of these stereotypes for the US and India, the African continent has been largely ignored. The AI is "blind" to the real, complex, and nuanced way people actually think about each other in places like Nigeria or South Africa.
The Solution: The SAFARI Project
The authors created a new dataset called SAFARI (Sub-Saharan African Associations about Regionally-salient Identities).
Instead of just asking people to fill out a form on a computer (which is hard when internet is slow or keyboards don't have the right letters for local languages), they did something different. They acted like radio hosts.
- The Phone Call Method: They hired local partners to call people on the phone.
- Speaking the Language: The callers spoke in the person's native language (like Swahili, Yoruba, or Zulu), not just English. This is crucial because some ideas simply don't translate well, and people express themselves better in their mother tongue.
- The "Oral Tradition": Africa has a strong tradition of storytelling and oral history. By using phones, the researchers respected this tradition, capturing the tone and feeling of the stereotypes, not just the words.
What Did They Find?
They talked to 410 people from four countries and collected over 3,500 stereotypes in English and 3,200 in 15 different local languages.
Here are some interesting patterns they found, like discovering hidden treasures in a map:
- Magic and Rituals: In Ghana and Kenya, many stereotypes revolve around "witchcraft" or "rituals," often targeting specific ethnic groups.
- Money Matters: In Kenya and Nigeria, there are many stereotypes about who is "rich" or "showing off wealth."
- Jobs: In Ghana, people have strong opinions about certain professions, like police officers or politicians being corrupt.
- Gender: About 58% of the gender stereotypes targeted women.
The "Translation" Challenge
There was a tricky part. The researchers had to record the answers in English first (because their computers couldn't type in all those local languages easily) and then hire experts to translate them back into the local languages.
Think of it like this: You ask a friend a question in their native tongue, write down their answer in English, and then hire a translator to write it back in their native tongue. You might lose a little bit of the "flavor" or the specific emotion, but it's the only way to get the data into a format computers can understand. The team was very careful to keep the original meaning as intact as possible.
Why Does This Matter?
The researchers tested this new dataset against the world's most famous AI models (like the ones from Google, OpenAI, and Anthropic).
The result was a wake-up call.
Even the smartest AIs were still repeating these harmful stereotypes. In fact, the AI was often more likely to repeat the stereotype when speaking English than when speaking the local language. This is because the AI's "training data" is so heavy on Western English that it doesn't understand the local context.
The Big Takeaway
This paper isn't just about a list of words; it's about a new way of doing research.
- Old Way: "We will send a survey link to everyone." (Fails if you don't have good internet or a good keyboard).
- SAFARI Way: "We will call you, speak your language, and listen to your stories."
By building this "community-engaged" approach, the authors are saying: If you want AI to be safe and fair for everyone, you can't just build it in a Silicon Valley lab. You have to go to the community, listen to them, and build the tools with them.
This dataset is now a tool for developers to "teach" their AI models to stop repeating these harmful rumors and start understanding the real, diverse, and complex human stories of Sub-Saharan Africa.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.