← Latest papers
💬 NLP

Camellia: Benchmarking Cultural Biases in LLMs for Asian Languages

This paper introduces Camellia, a benchmark comprising over 19,000 annotated entities and 2,000 masked contexts across nine Asian languages, to evaluate and reveal significant cultural biases and context understanding gaps in multilingual Large Language Models regarding non-Western cultures.

Original authors: Tarek Naous, Anagha Savit, Carlos Rafael Catalan, Geyang Guo, Jaehyeok Lee, Kyungdon Lee, Lheane Marie Dizon, Mengyu Ye, Neel Kothari, Sahajpreet Singh, Sarah Masud, Tanish Patwa, Trung Thanh Tran, Zo
Published 2026-05-28
📖 5 min read🧠 Deep dive

Original authors: Tarek Naous, Anagha Savit, Carlos Rafael Catalan, Geyang Guo, Jaehyeok Lee, Kyungdon Lee, Lheane Marie Dizon, Mengyu Ye, Neel Kothari, Sahajpreet Singh, Sarah Masud, Tanish Patwa, Trung Thanh Tran, Zohaib Khan, Alan Ritter, Tanmoy Chakraborty, Yuki Arase, Keisuke Sakaguchi, JinYeong Bak, Wei Xu

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have a group of very smart, well-traveled robots (Large Language Models, or LLMs). These robots have read almost everything on the internet and can speak many languages. You might think that because they've read so much, they understand every culture equally well. But this paper, called Camellia, suggests that these robots are actually quite biased, especially when it comes to Asian cultures.

Here is a simple breakdown of what the researchers did and what they found, using some everyday analogies.

The Problem: The "Western Lens"

Think of these AI robots as tourists who have spent most of their time in New York and London. They know those places inside out. But when you ask them about a village in rural India or a neighborhood in Seoul, they sometimes get confused. They might assume that what works in New York is the "default" or "correct" answer, even when the context is clearly Asian.

The researchers wanted to know: Do these robots treat Asian names, foods, and places with the same respect and understanding as Western ones?

The Tool: The "Camellia" Test

To find out, the team built a giant test called Camellia. Think of Camellia as a massive, multi-lingual "spot the difference" game.

  • The Players: They tested four popular AI models (Llama, Qwen, Aya, and Gemma).
  • The Languages: They focused on 9 Asian languages (like Chinese, Japanese, Korean, Hindi, Urdu, etc.) covering 6 distinct cultures.
  • The Data: They created two main things:
    1. A List of 19,530 Items: This includes names, foods, drinks, cities, and sports teams. They made sure to have a "Western" version (like "Lasagna" or "New York") and an "Asian" version (like "Biryani" or "Seoul") for every language.
    2. 2,173 "Fill-in-the-Blank" Stories: They took real sentences from social media and hid the important word (the entity) behind a [MASK] token.
      • Example: "When I go to Korea, I absolutely have to try [MASK]." (The answer should be a Korean dish, not a pizza).

The Three Games They Played

The researchers asked the AI to play three different games to see how biased it was:

1. The "Cultural Fit" Game (Context Adaptation)

  • The Setup: The AI sees a sentence about a specific culture (e.g., a story about a Korean festival) and has to guess the missing word.
  • The Test: Does the AI guess a Korean name/food, or does it accidentally guess a Western one?
  • The Result: The AI struggled. In about 30% to 40% of cases, the robot guessed a Western item even when the story was clearly about an Asian culture. It's like a tourist in Tokyo trying to order sushi but accidentally asking for a hamburger because they think "burger" is the default food.

2. The "Mood Ring" Game (Sentiment Association)

  • The Setup: The AI reads a sentence with a hidden word and has to guess if the sentence is happy, sad, or neutral.
  • The Test: Does the AI think sentences with Western names are more negative, or sentences with Asian names are more positive?
  • The Result: The AI's "mood" changed depending on who was in the story. Some models were more likely to think Western entities were "bad" or "negative," while others thought Asian entities were "good." It's like a person who subconsciously feels a story is sadder just because the main character has a foreign name.

3. The "Treasure Hunt" Game (Extractive QA)

  • The Setup: The AI reads a long paragraph and has to find a specific name or place hidden inside it.
  • The Test: Can it find the Asian name as easily as the Western name?
  • The Result: The AI was much better at finding Western names than Asian ones. In some languages, the accuracy gap was huge (up to 20%). It's like a detective who can easily spot a famous Western celebrity in a crowd but misses the local hero standing right next to them.

Why Does This Happen?

The paper suggests a few reasons, which can be compared to how a library is organized:

  • The "Library" Problem: The AI was trained on data from the internet. If the internet has more books about Western culture than Asian culture, the AI learns more about the West.
  • The "Translator" Problem: The AI uses a "tokenizer" (a tool that breaks words into pieces). For some Asian languages, the AI's tokenizer is like a pair of glasses that is slightly blurry or missing pieces. It doesn't "see" the Asian words as clearly as the Western ones, making it harder to understand the context.
  • The "Origin" Problem: Models built in China (like Qwen) did better on Chinese, Japanese, and Korean tasks than models built elsewhere. This suggests that where the robot was "born" (trained) matters a lot.

The Bottom Line

The paper concludes that even though these AI models are "multilingual," they aren't "multicultural" yet. They often default to Western ideas, struggle to understand the nuances of Asian contexts, and treat Asian entities differently than Western ones.

The researchers built Camellia not to fix the robots immediately, but to give them a report card so we can see exactly where they are failing. They hope this will help future developers build AI that treats all cultures with equal fairness and understanding.

Important Note: The paper does not claim these robots are dangerous in a clinical sense or that they should be used for specific real-world decisions right now. It simply says: "Here is a test, and here is proof that the robots are biased. We need to fix this before we trust them with important tasks."

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →