← Latest papers
🤖 AI

Safety Alignment Illusion: The Cross-Lingual Safety Gap in LLMs

This paper introduces INCLUDE, a multilingual benchmark evaluating Indian-centric socio-cultural biases across six languages, revealing that current English-centric safety alignment in LLMs creates a critical cross-lingual safety gap where non-English outputs, particularly in Bengali and closed-source models, exhibit significantly higher bias than their English counterparts.

Original authors: Namya Bhatnagar

Published 2026-08-20
📖 5 min read🧠 Deep dive

Original authors: Namya Bhatnagar

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine a world where the voice of a computer assistant is as fluent and fair in a rural village in India as it is in a bustling office in London. For years, the technology behind these assistants has been trained primarily on English text, learning to recognize harmful stereotypes and refuse offensive requests through a process called safety alignment. This training acts like a filter, designed to stop the machine from repeating biases about race, religion, or social status. However, a critical question has remained unanswered: does this filter work just as well when the user speaks Hindi, Bengali, or a mix of English and Hindi? If the filter only understands English, then millions of people speaking other languages are left exposed to the very biases the technology was meant to prevent. This is the core concern of a new study that investigates whether the promise of safe artificial intelligence holds true across the diverse linguistic landscape of India.

The researchers, led by Namya Bhatnagar, set out to measure this gap by creating a specialized test called INCLUDE. This benchmark was designed to see how well different artificial intelligence models handle culturally specific stereotypes when prompted in six different languages: English, Hindi, Bengali, Marathi, Tamil, and Hinglish, which is a common mix of Hindi and English used in daily conversation. To build this test, the team first defined six areas where bias often appears, such as caste, religion, and region. They then gathered a panel of twenty experts, including teachers and university faculty, to review and refine thousands of questions. These questions were carefully crafted to sound natural, asking the models to complete sentences in ways that would reveal if they were leaning toward harmful stereotypes or rejecting them. The team translated these prompts into the five Indian languages, ensuring that the meaning and the intensity of the stereotypes remained consistent across all versions.

The study evaluated ten different artificial intelligence models, splitting them into two groups: those that are open for anyone to inspect and modify, and those that are closed systems owned by large technology companies. When the researchers tested the open-source models, they found a clear and troubling pattern. The models were significantly more likely to produce biased, stereotypical answers when the prompts were written in Indian languages compared to English. Among the Indian languages, Bengali triggered the highest level of bias, followed closely by Hindi. Surprisingly, English prompts actually resulted in the lowest bias scores for these open models, suggesting that the safety filters are working well for English speakers but failing for everyone else. This creates an illusion of safety where a user might feel protected in English but is left vulnerable in their native tongue.

The results for the closed-source models, which are often used to power the search and retrieval functions of voice assistants, told a different and even more complex story. In this group, the pattern flipped completely. The models showed the strongest bias associations when the prompts were in English, while the Indian languages showed lower bias scores. This reversal suggests that the source of the problem is not just one thing. For the open models, the issue appears to be that the safety training was done mostly in English, leaving other languages unguarded. For the closed models, the bias seems to come from the massive amount of English text they were originally trained on, which has embedded strong stereotypes into their core understanding of the world.

One of the most significant findings concerned Hinglish, the code-mixed language used by millions of people in India. The study found that Hinglish did not get the safety benefits of English. Even though it contains English words, the models treated it statistically the same as the other low-resource Indian languages, grouping it with Hindi and Tamil rather than with English. This means that users who rely on this mixed language for voice interaction are not receiving the same level of protection as those speaking pure English. The researchers also noted that while the bias scores varied, the differences were not random; they were consistent enough to be considered a real, measurable failure in how these systems are built.

The study concludes that the safety of artificial intelligence is not uniform across languages. It reveals that the current methods of training these models create a divide where English speakers are protected from harmful stereotypes, while speakers of other languages are not. The findings suggest that simply translating safety rules from English is not enough; the very structure of how these models learn and how they are aligned needs to change to account for the nuances of different cultures and languages. By identifying these gaps, the researchers hope to provide a roadmap for developers to build systems that are truly fair for everyone, regardless of the language they speak. The work highlights that achieving true fairness in artificial intelligence requires looking beyond the dominant languages and addressing the specific cultural contexts in which these technologies are deployed.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →