← Latest papers
💬 NLP

'Ghaib in Translation' aka Unseen Harm: Measuring Cross-Script Safety Inconsistency with 'Missed-in-Urdu' Scores in LLM Hate Speech Detection

This paper reveals that large language models exhibit significant safety inconsistencies and "Missed-in-Urdu" rates when detecting hate speech in Urdu across various scripts, highlighting a critical gap in content moderation reliability caused by the complete absence of Urdu in nine years of major AI safety research proceedings.

Original authors: Fawzia Zehra (Fuzzy), Kara-Isitt, Sonal Khosla, Stephen Swift

Published 2026-08-26
📖 5 min read🧠 Deep dive

Original authors: Fawzia Zehra (Fuzzy), Kara-Isitt, Sonal Khosla, Stephen Swift

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

The internet is a vast, noisy marketplace where billions of people share thoughts, jokes, and grievances every second. To keep this space safe, technology companies rely on automated systems that act as digital gatekeepers, scanning posts to identify hate speech and threats before they can cause real-world harm. These systems are trained on massive amounts of text, learning to recognize patterns of abuse in the languages they know best. However, the world speaks many languages, and the ability of these digital gatekeepers to understand them is not equal. Some languages, like English, are well-represented in the data these systems learn from, while others are often overlooked. When a system fails to recognize hate speech in a specific language, it creates a blind spot where harm can spread unchecked, even as the same system successfully blocks that same harm if it were written in a different tongue. This gap is not just a technical glitch; it is a safety failure that leaves millions of speakers vulnerable.

A recent study by researchers at Brunel University London and the Hasso Plattner Institute in Germany shines a light on one of the most significant of these blind spots: Urdu. Urdu is the tenth most spoken language in the world, with 246 million speakers across Pakistan, India, and large communities in the United Kingdom and the United States. Despite its size and the high volume of online abuse documented in Urdu-speaking communities, the researchers found that the field of online safety research has almost entirely ignored it. In a comprehensive review of nine years of major academic conferences dedicated to online abuse, the team discovered that not a single paper was dedicated to Urdu. This absence is striking, especially since a major workshop in 2022 had explicitly called for more work on the language. The researchers set out to see if this lack of attention had measurable consequences for the safety tools currently in use.

To test this, the team examined five of the most advanced artificial intelligence models available today. They fed these models a large collection of text containing hate speech, threats, and normal conversation, but they presented the content in different ways. Some text was in the original Urdu script, known as Nastaliq, which flows from right to left. Other text was written in Roman Urdu, where Urdu words are spelled out using English letters, a common practice on social media. The researchers also provided English translations of the same messages. By comparing how the models reacted to the same message in these different formats, they could see if the safety systems were truly understanding the content or if they were simply reacting to the shape of the letters.

The results revealed a troubling inconsistency. When the models read the harmful content in English, they correctly identified it as abusive. However, when the exact same message was presented in the original Urdu script, the models often failed to flag it as dangerous. On average, across the different models tested, about 18 percent of the harmful messages were labeled differently depending on whether they were in Urdu or English. In simpler terms, if a person posted a threat in Urdu, the system might let it pass as a normal post, but if that same threat were translated into English, the system would immediately block it. This "missed-in-Urdu" rate, where harmful content slips through the cracks only because of the script it is written in, ranged from roughly 2 percent to nearly 10 percent depending on the specific model used.

The study also highlighted that not all artificial intelligence models are equally vulnerable to this problem. The most advanced, large-scale models, often called frontier models, showed better consistency, with fewer errors. However, smaller, open-source models showed much higher rates of failure, missing nearly a third of the harmful content when it was in Urdu. This suggests that the safety features of these systems are not uniform; they are unevenly distributed, offering stronger protection to speakers of some languages while leaving others exposed. The researchers also noted a structural gap in the available data: there is no dedicated dataset for hate speech that mixes Urdu and English, a common way people communicate online. Without this specific type of data, it is impossible to fully test how well these systems handle the complex, mixed-language reality of modern social media.

Ultimately, the work demonstrates that current safety tools provide an uneven shield. A system that can reliably detect hate speech in English but fails to recognize the same speech in Urdu is not truly safe for the millions of people who speak that language. The findings confirm that the absence of Urdu from safety research is not just an academic oversight; it has real, measurable consequences for content moderation. Until these gaps are filled with better data and more rigorous testing across different scripts, a significant portion of the world's online population will remain vulnerable to harm that the very systems designed to protect them simply cannot see.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →