← Latest papers
💬 NLP

The Masked Advantage: Uncovering Local-Language Access to Cultural Knowledge in LLMs

This paper reveals that while large language models often exhibit higher raw accuracy in English due to superior general proficiency, their access to local cultural knowledge is frequently more effective in native languages once this proficiency gap is controlled for, suggesting that lower local-language performance often masks, rather than reflects, weaker cultural understanding.

Original authors: Yang Zhang, Xiao Fei, Amr Mohamed, Sarah Almeida Carneiro, Mersin Konomi, Mingmeng Geng, Ahmed Asaad, Guokan Shang, Michalis Vazirgiannis

Published 2026-06-08
📖 5 min read🧠 Deep dive

Original authors: Yang Zhang, Xiao Fei, Amr Mohamed, Sarah Almeida Carneiro, Mersin Konomi, Mingmeng Geng, Ahmed Asaad, Guokan Shang, Michalis Vazirgiannis

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have a library of books about the history, traditions, and daily life of 13 different countries. You also have a team of 80 different "super-readers" (AI models) who can read these books and answer questions about them.

The big question this paper asks is: To get the best answers about a specific country, should you ask the super-reader in English, or in that country's local language?

The Problem: The "Fluency Fog"

Previously, researchers looked at raw scores and got confused. They saw that when you asked an AI about French culture in French, it often got the answer wrong. When asked in English, it got it right. They concluded, "Aha! The AI knows more about French culture when we speak English!"

The authors of this paper say: Wait a minute. That's like judging a chef's cooking skills by how well they can read the recipe in a foreign language. If the chef is a master of French cuisine but can't read English well, they might fail a test written in English, not because they don't know the food, but because they can't read the instructions.

The paper argues that AI models are often fluent in English but clumsy in local languages. This "fluency fog" hides the fact that the AI might actually know the local culture better when you speak to it in its own tongue.

The Solution: The "Three-Step Detective"

To solve this, the researchers built a clever testing framework using a mathematical tool called IRT (Item Response Theory). Think of IRT as a special scale that weighs the difficulty of the question against the skill of the reader, so you can compare apples to apples even if the questions are different.

They tested the AI with three types of comparisons:

  1. The "Global" Test (Proficiency Check):

    • The Setup: Ask the AI general questions (like "What is the capital of France?") in both English and French.
    • The Result: The AI almost always does better in English.
    • The Meaning: This measures Language Proficiency. It tells us the AI is just better at reading and understanding English. This is the "fluency fog."
  2. The "Local" Test (Raw Performance):

    • The Setup: Ask the AI specific cultural questions (like "What is the traditional dish for a specific local festival?") in both English and French.
    • The Result: Mixed. Sometimes English wins, sometimes French wins.
    • The Meaning: This is the messy real-world score. It's a mix of "How good is the AI at French?" + "How much French culture does it actually know?"
  3. The "Knowledge" Test (The Magic Step):

    • The Setup: The researchers took the "Local" score and subtracted the "Global" score.
    • The Logic: If you remove the "Language Proficiency" part, what is left? The pure Cultural Knowledge Access.
    • The Result: Boom! In almost every single case, the AI knew the local culture better when asked in the local language.

The Big Discovery: The "Masked Advantage"

The paper calls this the "Masked Advantage."

Imagine a brilliant local historian who speaks perfect English but has a slight stutter when speaking their native tongue.

  • If you ask them a history question in English, they answer perfectly.
  • If you ask them in their native tongue, they stumble over the words and give a wrong answer, even though the facts are in their head.

The paper found that for AI, the "stutter" (weak language skills) is so strong that it masks the "facts" (cultural knowledge). The AI actually holds the local cultural knowledge more accessibly in the local language, but because it struggles to form the sentences, it looks like it knows less.

Key Takeaways from the Study

  • English is the "Safe Bet" for general smarts: The AI is consistently better at English. If you just want to know general facts, English is the most reliable language.
  • Local languages hold the "Secret Knowledge": Once you account for the AI's struggle with grammar and vocabulary, the local language is actually the key that unlocks the cultural knowledge.
  • Bigger isn't always better (for this specific issue): Even the newest, biggest AI models show this pattern. They know the culture better in the local language, but their English fluency is just so much higher that it overshadows the local advantage.
  • Regionally trained models help: AI models that were specifically trained on data from a specific region (like a model trained mostly on Arabic data) show a stronger local advantage, but the "masked" effect still exists.

The Bottom Line

The paper concludes that weaker performance in a local language doesn't mean the AI lacks cultural knowledge. It usually just means the AI is bad at speaking that language. The knowledge is there, waiting to be accessed, but it's currently hidden behind a wall of language difficulty.

To get the truest picture of a culture from an AI, you shouldn't just look at the final score; you have to look past the language barrier to see what the AI actually knows.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →