← Latest papers
💬 NLP

Who Pays More for Safety? Measuring the Disparate Cost of Safety Alignment across Languages

This paper introduces a "Safety Cost" metric to demonstrate that safety alignment disproportionately penalizes non-English languages with greater utility loss and weaker protection compared to English, revealing a systematic inequity in current safety practices.

Original authors: Chanwoong Yoon, Jungsoo Park, Alan Ritter

Published 2026-08-25
📖 4 min read☕ Coffee break read

Original authors: Chanwoong Yoon, Jungsoo Park, Alan Ritter

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

In the world of artificial intelligence, large language models are the digital engines that power chatbots, translators, and assistants. Before these machines can be useful to people, they undergo a rigorous training process called alignment. Think of this as a period of schooling where the model learns not just to answer questions, but to answer them safely, adhering to human values and refusing to generate harmful content like instructions for violence or illegal acts. This safety training is essential, but it comes with a hidden price. Just as a car with overly sensitive brakes might stop too quickly for a minor obstacle, an AI that is too cautious can become unhelpful, refusing to answer harmless questions or providing vague, watered-down responses. This loss of helpfulness is known as a utility cost. For years, researchers have known that safety training reduces a model's overall performance, but a critical question remained unanswered: does this cost fall equally on everyone, or do some languages suffer more than others?

A team of researchers set out to investigate whether the burden of safety alignment is shared fairly across the world's languages. They focused on a specific phenomenon where models might reject safe requests simply because they sound slightly risky, a behavior known as over-refusal. To measure the true cost of safety, the team devised a clever comparison method. They took powerful, safety-aligned models and created "unaligned" versions of the same software. These unaligned versions were not trained to be dangerous; rather, the researchers surgically removed the specific safety mechanisms that caused the models to refuse requests, leaving the rest of the model's knowledge and ability to speak intact. By comparing the responses of the safety-aligned model against its unaligned twin for the same question, the researchers could isolate exactly how much helpfulness was lost due to safety training alone. They tested this across English and five other languages, including Chinese, Arabic, Korean, Vietnamese, and Thai, using a set of questions designed to be harmless but likely to trigger a safety alarm.

The results revealed a stark and systematic inequality. The researchers found that non-English users consistently pay a higher price for safety than English speakers. In many cases, the safety filters were less effective at protecting non-English speakers from actual harm, yet these users still received significantly worse, less helpful answers to their harmless questions. The study identified three distinct patterns behind this disparity. First, many languages fell into a "double penalty" zone: they received weaker protection against real dangers while simultaneously suffering a much larger drop in the quality of their answers. Second, in some instances, a language appeared to have a low cost of safety only because the safety filters had failed to engage at all, leaving users exposed to harmful content without the model even realizing it. Third, even for high-resource languages like Chinese, which have vast amounts of training data, the cost of achieving the same level of safety as English was substantially higher. The utility loss was not just about the model saying "no" when it should have said "yes." The researchers discovered that even when the model did answer, the response was often thinner, less detailed, and factually less precise. The safety training had quietly stripped away depth and nuance, making the AI less useful without the user necessarily realizing why.

This work suggests that the current methods for making AI safe are not calibrated fairly across different cultures and languages. The safety training, which is heavily influenced by English-language data and norms, inadvertently creates a structural gap where non-English speakers receive a poorer experience. The researchers argue that this is not an unavoidable side effect of making AI safe, but rather a flaw in how safety is currently implemented. By measuring the specific cost of safety in each language, the study exposes a fundamental inequity in the technology. It shows that while English speakers might get a safe and helpful response, a speaker of another language might get a response that is either unsafe or unhelpfully vague, all because the safety mechanisms are not tuned to their specific linguistic context. The findings call for a new approach to safety alignment, one that ensures the cost of being safe is distributed evenly, so that the benefits of artificial intelligence are not diminished by the language in which it is spoken.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →