← Latest papers
💻 computer science

LLM Safety Alignment in Low-Resource Languages: A Systematic Literature Review

This systematic literature review analyzes 50 studies on LLM safety alignment in low-resource languages, identifying a persistent multilingual safety gap driven by factors like uneven pre-training and culturally blind benchmarks, while proposing a new taxonomy and calling for culturally grounded solutions to address vulnerabilities such as cross-lingual jailbreaks and safety degradation.

Original authors: Valdini Douglace Lemofouet, Blessing Ngozi Uzor, Paula Chikaodinaka Anyanwu, Danielle Blanche Kapsa, Sukairaj Hafiz Imam, P Sam Sahil, Abigail Oppong, Tassallah Abdullahi, Clemencia Siro, Idris Abdulm
Published 2026-08-18
📖 5 min read🧠 Deep dive

Original authors: Valdini Douglace Lemofouet, Blessing Ngozi Uzor, Paula Chikaodinaka Anyanwu, Danielle Blanche Kapsa, Sukairaj Hafiz Imam, P Sam Sahil, Abigail Oppong, Tassallah Abdullahi, Clemencia Siro, Idris Abdulmumin, Seid Muhie Yimam, Shamsuddeen Hassan Muhammad

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine a world where the most advanced artificial intelligence systems, capable of writing poetry, diagnosing illnesses, and answering complex questions, suddenly become unreliable or even dangerous the moment you switch the language from English to something else. This is not a hypothetical scenario but a growing reality in the field of artificial intelligence. Large language models are computer programs trained on vast amounts of text to understand and generate human language. To make these tools safe for public use, researchers teach them to refuse harmful requests, such as instructions to build weapons or spread hate speech. This teaching process is known as safety alignment. For years, this training has happened almost exclusively using English data. As a result, these digital minds have learned to be polite and cautious in English, but they often fail to recognize the same dangers when a user speaks in a language with fewer digital resources, such as many African or indigenous languages. The gap between how safe these models are in English versus how safe they are elsewhere is widening, creating a situation where speakers of underrepresented languages face higher risks of receiving toxic, biased, or misleading advice from the very tools meant to help them.

A team of researchers set out to map this uneven landscape by conducting a systematic review of fifty recent studies focused on safety in low-resource languages. They gathered their findings from thousands of academic papers, filtering them down to the most relevant work to understand exactly where the technology is failing and why. Their investigation revealed a stark truth: the methods used to keep AI safe in English do not simply translate to other languages. When researchers tried to use English safety rules for languages like Hausa, Swahili, or Yoruba, the systems often broke down. The study found that simply translating a harmful English prompt into a low-resource language could trick the AI into ignoring its safety guardrails nearly eighty percent of the time, whereas the same trick in English would fail most of the time. This suggests that the safety lessons learned in English are not deeply embedded in the model's understanding of the world but are instead tied too closely to the English language itself.

The researchers identified three main ways scientists are trying to fix this problem. The first approach involves creating new training data that is rooted in local cultures rather than just translating English examples. One study showed that when they built a safety dataset specifically for local languages, using examples of harm that actually exist in those communities, the AI became significantly better at refusing dangerous requests. In contrast, models trained on translated English data often missed the nuance of local cultural harms. The second approach tries to teach the model to be safe in one language and then hope that knowledge spreads to others. However, the review suggests this is a fragile strategy. When models are fine-tuned on new languages, even with harmless data, they sometimes forget the safety rules they learned in English. The third approach looks inside the model's own structure, attempting to tweak specific internal settings to make it safer without needing massive amounts of new data. While promising, these methods are still in their early stages and have not yet been proven to work reliably across all languages.

A major part of the problem lies in how we test these systems. Most safety tests are built by translating English questions into other languages. The researchers found that this method is flawed because it fails to capture the specific ways harm manifests in different cultures. For instance, a model might refuse a request to generate hate speech in English but might happily generate toxic product recommendations in Hausa because the concept of that specific harm was not part of its training. The review highlighted that there are very few safety benchmarks designed specifically for African languages, leaving a huge blind spot in our understanding of how these models behave. Without tests that reflect real-world cultural contexts, developers cannot know if their safety measures are actually working.

The study also uncovered a surprising vulnerability: the act of teaching a model a new language can sometimes make it less safe overall. When researchers added new languages to a model that was already safe, the model's ability to refuse harmful requests in its original languages sometimes degraded. This indicates that the process of learning new linguistic patterns can interfere with the safety constraints already in place. Furthermore, the review noted that when people mix languages in a single sentence, a common practice in many parts of the world, the AI becomes much easier to trick. These mixed-language prompts can bypass safety filters that would stop a pure-language attack.

Ultimately, the review concludes that the current path of simply translating English safety data is insufficient. The researchers argue that true safety requires a shift toward building datasets and evaluation tools from the ground up in local languages, involving native speakers to define what constitutes harm in their own cultural contexts. They suggest that future progress depends on balancing the training data so that no single language dominates the model's understanding of safety. Until these steps are taken, the promise of safe artificial intelligence will remain out of reach for billions of people who speak languages other than English, leaving them exposed to risks that their English-speaking counterparts do not face.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →