← Latest papers
💬 NLP

Information Asymmetry across Language Varieties: A Case Study on Cantonese-Mandarin and Bavarian-German QA

This paper introduces a novel QA dataset to demonstrate that Large Language Models struggle to answer questions about knowledge unique to local language varieties (Cantonese and Bavarian) compared to their standard counterparts (Mandarin and German), highlighting significant information asymmetries and the critical need for improved cultural inclusivity in AI.

Original authors: Renhao Pei, Siyao Peng, Verena Blaschke, Robert Litschko, Barbara Plank

Published 2026-03-17
📖 4 min read☕ Coffee break read

Original authors: Renhao Pei, Siyao Peng, Verena Blaschke, Robert Litschko, Barbara Plank

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have a giant, all-knowing librarian named LLM (Large Language Model). This librarian has read almost every book in the world, but there's a catch: they mostly read the "Standard Edition" of books. They know everything about the "Official Version" of a topic, but they often miss the colorful, local details found in the "Regional Edition."

This paper is like a detective story where the authors ask: "What happens when you ask our librarian about a local secret that only exists in the Regional Edition?"

Here is the breakdown of their investigation using simple analogies.

1. The Problem: The "Local Secret" Gap

Think of Wikipedia as a massive library with different sections for different languages.

  • The Standard Section: This is like the main library in the city center. It's huge, well-organized, and has millions of books (e.g., Mandarin Chinese, Standard German).
  • The Local Section: This is like a cozy, neighborhood library. It has fewer books, but it holds local secrets that the big city library doesn't know about.
    • Example: A specific folk dance in Bavaria might have a detailed history in the Bavarian dialect Wikipedia, but the Standard German Wikipedia page just says, "It's a dance."
    • Example: A specific type of ancient Chinese opera might be described in Cantonese Wikipedia with details about its origins, while the Mandarin version skips those details.

The authors call this "Information Asymmetry." It's like the librarian knows the general story but missed the juicy local gossip.

2. The Experiment: The "Quiz Show"

The researchers built a special quiz show called WILOVA-QA.

  • The Players: They tested several super-smart AI librarians (like Llama, Qwen, and GPT).
  • The Questions: They asked questions based only on the local secrets found in the neighborhood library (Cantonese and Bavarian pages) that were missing from the city library (Mandarin and German pages).

The Results (Round 1: Closed Book):
When they asked the AI, "Tell me this local secret" without giving them any help, the AI failed miserably.

  • Analogy: It's like asking a chef, "What's the secret ingredient in this specific family recipe?" and the chef says, "I don't know, I've never heard of it." Even though the recipe exists in the world, the chef's internal memory didn't have it.

The Results (Round 2: Open Book):
Then, the researchers gave the AI a cheat sheet. They handed them the text from the local Wikipedia page.

  • Scenario A (Local Language): They gave the text in the local dialect (e.g., Bavarian).
    • Result: The AI suddenly got it right! It could read the local text and answer the question.
  • Scenario B (Translated): They translated the local text into the standard language (e.g., translating Bavarian to German) and gave it to the AI.
    • Result: The AI did even better. It seems the AI is like a student who understands the lesson best when it's explained in the language they are most comfortable with, even if the original source was in a dialect.

3. The Big Discovery: The "Missing Pages"

The study found that these "local secrets" aren't just rare; they are systematically missing from the AI's brain.

  • If a fact is only in the local Wikipedia, the AI doesn't know it exists.
  • If you give the AI the local text, it can reason through it perfectly.
  • The Takeaway: The AI isn't "dumb"; it just hasn't been fed the "Regional Edition" of the encyclopedia during its training. It's like a student who studied the textbook but never read the local newspaper.

4. Why This Matters

The authors argue that this is a problem of inclusivity.

  • If we only train AI on "Standard" languages, we lose the unique cultural flavor, history, and local knowledge of millions of people.
  • The AI becomes a "Global Generalist" that knows a little bit about everything, but misses the deep, specific details that make a culture unique.

The Bottom Line

Imagine the AI as a tour guide.

  • Right now, if you ask the guide about a famous landmark, they give a perfect, standard tour.
  • But if you ask about a hidden, local street festival that only the locals know about, the guide says, "I don't know."
  • The Solution: If you hand the guide a local map (the context), they can instantly tell you everything about the festival.

The paper's conclusion: To make AI truly smart and culturally aware, we need to stop ignoring the "neighborhood libraries" and start feeding the AI the local, dialect, and regional stories that are currently missing from its memory.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →