← Latest papers
💬 NLP

Lost in Localization: Building RabakBench with Human-in-the-Loop Validation to Measure Multilingual Safety Gaps

This paper introduces RabakBench, a human-in-the-loop validated multilingual safety benchmark for Singapore's linguistic landscape that reveals significant performance gaps in state-of-the-art LLM guardrails when handling low-resource languages like Singlish, Chinese, Malay, and Tamil.

Original authors: Gabriel Chua, Leanne Tan, Ziyu Ge, Roy Ka-Wei Lee

Published 2026-02-03
📖 5 min read🧠 Deep dive

Original authors: Gabriel Chua, Leanne Tan, Ziyu Ge, Roy Ka-Wei Lee

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have a very smart, well-trained security guard for a building. This guard is excellent at spotting trouble when people speak in "Standard English." They know exactly what a rude comment, a threat, or a dangerous idea looks like in that language.

But now, imagine that same guard is asked to watch over a bustling, multicultural neighborhood where people speak a mix of languages, slang, and local dialects. Suddenly, the guard gets confused. They might miss a real threat because it's hidden in a local joke, or they might wrongly arrest a harmless neighbor because they don't understand a cultural phrase.

This paper, RABAKBENCH, is about building a new "test" to see how well these AI security guards handle that messy, real-world neighborhood. The neighborhood in question is Singapore, a place where people constantly mix English with Chinese, Malay, and Tamil words (a mix called "Singlish").

Here is how the researchers built their test and what they found, explained simply:

1. The Problem: The "Localization Blind Spot"

Current AI safety systems are like guards who only studied a textbook on Standard English. They fail when faced with:

  • Code-mixing: Sentences that jump between languages.
  • Slang: Local words that change meaning depending on context.
  • Dialects: Regional ways of speaking that aren't "proper" English.

The researchers found that these AI guards often make two big mistakes:

  1. False Negatives: They miss actual hate speech or threats because the words look like harmless slang to them.
  2. False Positives: They flag harmless cultural jokes as dangerous, effectively silencing normal conversation.

2. The Solution: Building "RABAKBENCH"

To fix this, the team built a massive testing ground called RABAKBENCH (named after a local word meaning "extreme" or "intense"). Think of it as a simulator for AI safety.

They didn't just write these test sentences themselves; they used a clever three-step factory process:

  • Step 1: The "Red Team" Attack (Generate)
    Imagine a group of hackers trying to trick the AI guards. The researchers used AI to generate thousands of tricky, real-world examples of Singlish, including things that sound dangerous but might be safe, and things that sound safe but are actually toxic. They even used a "critic" AI to help the "attacker" AI get better at finding weak spots.
  • Step 2: The "Super-Labelers" (Label)
    Labeling these tricky sentences by hand is expensive and hard. So, the researchers tested several different AI models to see which ones were best at understanding human judgment. They found three "star" AI models that agreed with human experts about 70–80% of the time. They used these AI models to label the data, but they kept a human in the loop to double-check the work, ensuring the cultural nuances were correct.
  • Step 3: The "Cultural Translator" (Translate)
    They took their Singlish test cases and translated them into Chinese, Malay, and Tamil. But here's the catch: standard translators often "sanitize" bad words to make them polite. The researchers built a special translation process that ensured the toxicity stayed the same. If a sentence was mean in Singlish, the translation had to be equally mean in Malay or Tamil, preserving the original "flavor" and intent.

The result? A dataset of over 5,000 examples covering four languages, with detailed labels on exactly why something is unsafe (e.g., is it hate speech? is it just an insult? is it a threat?).

3. The Results: The Guards Failed the Test

The researchers took 13 of the world's top AI safety systems (both commercial ones like Google and OpenAI, and open-source ones) and ran them through this new test.

The verdict was harsh:

  • Performance Dropped: Systems that scored 80%+ on standard English tests often dropped to below 50% on this local test.
  • The "Tamil Gap": The performance was especially poor for Tamil, where many systems scored below 30%.
  • One-Size-Fits-All Doesn't Work: The study showed that you can't just train a guard on English and expect them to work everywhere. They need specific training for local cultures.

4. Why This Matters

The paper argues that if we want AI to be safe for everyone, we can't just rely on standard English rules. We need to build safety systems that understand the "local flavor."

The Takeaway:
RABAKBENCH is like a driver's license test for AI safety in multicultural environments. It proves that current AI drivers are great on the highway (Standard English) but terrible in the city streets (local dialects). The researchers have released their test questions and answers to the public so that other developers can build better, more culturally aware safety guards for the future.

What the paper does NOT claim:

  • It does not claim to have fixed the AI systems yet; it only built the test to show they are broken.
  • It does not suggest these tests should be used for clinical or medical safety.
  • It does not promise that future AI will automatically be perfect; it just provides a roadmap for how to measure and improve them.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →