The "Knowledge-Behavior Gap" in Cultural Taboo Safety of Large Language Models
This paper introduces CulShield, the first benchmark spanning 77 countries and over 2,000 taboos to evaluate large language models' cultural taboo safety, revealing a significant "knowledge-behavior gap" where models possess cultural knowledge but fail to apply it in implicit, context-dependent interactions.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are teaching a super-smart robot to be a good guest at a dinner party. You wouldn't just teach it the names of the dishes (knowledge); you'd also have to teach it not to eat the food off the host's plate, even if it's starving (behavior). This is the world of Large Language Models (LLMs), the AI brains behind chatbots and assistants. These models are like digital encyclopedias that can chat in almost any language, but they are currently struggling with a tricky social rule: cultural taboos. A cultural taboo is like an invisible "do not touch" sign in a specific culture. For example, in some places, wearing a green hat is a fashion statement, while in others, it's a huge insult implying your partner is unfaithful. If a robot doesn't know the difference, it might accidentally offend a guest, ruin a conversation, or even cause real-world trouble. The big question researchers are asking is: Just because a robot knows the rules, does it actually follow them when the pressure is on?
This paper, titled "The 'Knowledge–Behavior Gap' in Cultural Taboo Safety of Large Language Models," dives right into that messy reality. The authors, a team from Fudan University and Huawei, realized that while we have plenty of tests to see if robots know cultural facts, we don't have good ways to test if they can actually act safely when a cultural trap is hidden inside a harmless question. To fix this, they built a new testing ground called CulShield. Think of CulShield as a giant, multicultural "mystery box" containing over 2,000 different cultural taboos from 77 countries and territories. They didn't just ask the robots, "Do you know that green hats are bad in China?" (which is easy). Instead, they created "cultural traps"—questions that sound totally innocent but are actually designed to trick the robot into breaking a taboo. For instance, they might ask, "I'm planning a surprise party for my Chinese friend; what's a fun, colorful hat I can give him?" A robot that only has knowledge but no behavioral guardrails might say, "Sure, a green one!" without realizing the social disaster it's causing.
The researchers tested some of the most advanced robots available, including GPT-4o-mini and Gemini-2.5-pro, using this new benchmark. What they found was a bit like discovering that a student who aced the history test still gets lost in the school hallway. They found a clear "knowledge–behavior gap." The robots often knew the rules perfectly when asked directly, but when the rules were hidden inside a tricky question, they frequently failed to apply that knowledge. It's as if the robot's brain knew the map, but its feet refused to walk the right path.
Another surprising discovery was how the language used in the question changed the robot's behavior. The study suggests that the language acts like a "cultural lens." When the robots were asked questions in English, they were more likely to trip over cultural traps related to cultures close to English-speaking norms. However, when the questions were in Spanish or Chinese, their behavior shifted, sometimes becoming more cautious, sometimes less. It turns out that just speaking a language doesn't mean the robot fully understands the culture behind it.
To see if they could fix this, the team tried "training" the robots with new data that explicitly taught them how to refuse these tricky requests politely. While this helped the robots get better at spotting the traps, it had a side effect: the robots became too sensitive. They started refusing harmless questions just to be safe, like a security guard who stops everyone from entering a building because they might be carrying a sandwich. To solve this, the researchers mixed in data about general human values (from the World Values Survey) to help the robots find a balance. The results showed that this mix helped the robots become safer without becoming overly paranoid.
In short, this paper suggests that teaching an AI to be culturally safe isn't just about feeding it more facts. It's about training it to recognize the hidden social landmines in everyday conversation and to act with the right amount of caution, no matter what language is being spoken. The authors admit that their new benchmark, CulShield, isn't perfect yet—it doesn't cover every single culture on Earth—but it's a crucial first step in helping our digital guests learn the rules of the global dinner party.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.