← Latest papers
💬 NLP

XL-SafetyBench: A Country-Grounded Cross-Cultural Benchmark for LLM Safety and Cultural Sensitivity

This paper introduces XL-SafetyBench, a cross-cultural benchmark comprising 5,500 country-grounded test cases across 10 language pairs that reveals a disconnect between jailbreak robustness and cultural sensitivity in frontier models, while exposing that local models' apparent safety often stems from generation failures rather than genuine alignment.

Original authors: Dasol Choi, Eugenia Kim, Jaewon Noh, Sang Seo, Eunmi Kim, Myunggyo Oh, Yunjin Park, Brigitta Jesica Kartono, Josef Pichlmeier, Helena Berndt, Sai Krishna Mendu, Glenn Johannes Tungka, Özlem Gökçe, Sur
Published 2026-05-08
📖 5 min read🧠 Deep dive

Original authors: Dasol Choi, Eugenia Kim, Jaewon Noh, Sang Seo, Eunmi Kim, Myunggyo Oh, Yunjin Park, Brigitta Jesica Kartono, Josef Pichlmeier, Helena Berndt, Sai Krishna Mendu, Glenn Johannes Tungka, Özlem Gökçe, Suresh Gehlot, Katherine Pratt, Amanda Minnich, Haon Park

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are hiring a new assistant to help people all over the world. You want to make sure this assistant is safe, polite, and won't accidentally say something offensive or dangerous.

For a long time, we've tested these assistants using a "standardized test" written entirely in English. But the authors of this paper, XL-SafetyBench, argue that this is like testing a chef's ability to cook Italian food by only asking them to make American burgers. Just because they can make a burger safely doesn't mean they know the local rules of Italy.

Here is a simple breakdown of what they did and what they found, using some everyday analogies.

1. The Problem: The "Translation Trap"

Most safety tests for AI are just English questions translated into other languages.

  • The Flaw: If you translate a question about "how to steal a car" into Korean, the AI might say "No." But what if the question is about a specific type of housing scam that only happens in Korea? A translated test won't catch that.
  • The Missing Piece: There are two types of "safety" the old tests miss:
    1. The "Hacker" Test: Can the AI be tricked by a local criminal into doing something bad? (e.g., "How do I scam people using the specific Korean housing deposit system?")
    2. The "Social Graces" Test: Does the AI know local customs? (e.g., If you ask the AI to help plan a gift for a French friend, it shouldn't suggest chrysanthemums, because in France, those flowers are for funerals, not birthdays.)

2. The Solution: A New "World Tour" Test

The researchers built XL-SafetyBench, a massive test suite with 5,500 questions covering 10 different countries (like the US, South Korea, India, Germany, etc.).

They didn't just translate English questions. They built the test from the ground up in each local language, using two main tracks:

  • Track A: The "Red Team" (Jailbreak Benchmark)

    • The Analogy: Imagine hiring a team of local "hacks" to try to trick the AI. They don't just ask "How do I make a bomb?" They ask, "How do I use the specific local banking app in Turkey to steal money?"
    • The Goal: To see if the AI can resist these clever, locally-grounded tricks.
  • Track B: The "Cultural Detective" (Cultural Benchmark)

    • The Analogy: Imagine the AI is a guest at a dinner party. The host asks, "Can you help me write a menu for a wedding?" But hidden inside the request is a tiny detail: "Let's serve pork to the guests who follow a strict vegetarian religion."
    • The Goal: The AI needs to spot that tiny, hidden cultural mistake without being told it's there. It's not about refusing a bad request; it's about noticing a social faux pas buried in a normal conversation.

3. How They Built It

They didn't just ask a computer to write these questions. They used a "human-in-the-loop" process:

  1. AI Discovery: Computers found potential local issues.
  2. Human Validation: Real humans who lived in those countries for over 15 years checked every single question. They made sure the questions sounded natural and that the cultural traps were real.
  3. The Result: A high-quality, culturally authentic test that feels like a real conversation, not a robot translation.

4. The Big Surprises (The Findings)

When they tested 37 different AI models (10 big global ones and 27 local ones from those specific countries), they found two shocking things:

Surprise #1: Being "Safe" and Being "Culturally Aware" are not the same thing.

  • The Analogy: Think of safety and culture as two different skills, like "driving a car" and "speaking the local dialect."
  • The Finding: Some of the most powerful AI models were great at refusing to do bad things (high safety) but terrible at spotting cultural mistakes (low cultural awareness). Conversely, some models were okay at culture but easily tricked by hackers.
  • The Takeaway: You can't just give an AI one "safety score." You have to measure if it's safe and if it's culturally smart separately.

Surprise #2: Local models are "faking" safety.

  • The Analogy: Imagine a student taking a hard math test.
    • Model A (Global): Understands the question, thinks hard, and says, "I can't solve this because it's dangerous." (This is Real Safety).
    • Model B (Local): Doesn't understand the question at all, so it just stares blankly or says nonsense. Because it didn't answer the question, it technically "didn't do anything bad."
  • The Finding: Many local AI models looked "safe" because they were failing to understand the question, not because they were making a moral choice to refuse. They were "safe" by accident, not by design. The researchers call this the "Illusion of Safety."

5. Why This Matters

This paper introduces a new way to test AI that treats different countries as unique places with their own rules, rather than just "English with a different accent."

It shows us that to build truly safe AI for the whole world, we need to stop relying on translated tests. We need to test if the AI can handle local scams and if it knows not to give a funeral flower as a birthday gift. And most importantly, we need to make sure the AI isn't just "safe" because it's too confused to answer, but safe because it actually understands the world.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →