← Latest papers
💬 NLP

Attributing Culture-Conditioned Generations to Pretraining Corpora

This paper introduces the MEMOed framework to demonstrate that cultural biases in large language model generations stem from uneven pretraining data distributions, where high-frequency cultures trigger memorized outputs while low-frequency cultures often fail to generate relevant content.

Original authors: Huihan Li, Arnav Goel, Keyu He, Xiang Ren

Published 2026-04-21
📖 6 min read🧠 Deep dive

Original authors: Huihan Li, Arnav Goel, Keyu He, Xiang Ren

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

The Big Picture: The "Cultural Library" Problem

Imagine a massive library (the Pretraining Corpus) where a robot student (the AI Model) spends its entire childhood reading every single book, article, and website available.

The problem? This library is heavily biased. It has thousands of copies of books about American pizza, Italian fashion, and Japanese sushi. But for some cultures, like certain tribes in the Amazon or small island nations, the library might only have one or two thin pamphlets, or maybe just a single sentence mentioning them.

When you ask this robot, "What does a person from [Culture X] eat?" or "What do they wear?", the robot doesn't "know" the answer in the human sense. Instead, it tries to guess based on what it read most often.

This paper asks: Is the robot actually remembering the truth about that culture, or is it just guessing based on what it read the most?


The Detective Tool: MEMOED

The researchers built a detective tool called MEMOED (MEMOrization from pretraining document). Think of MEMOED as a forensic librarian.

When the robot gives an answer (e.g., "A person from India eats Biryani"), MEMOED goes back to the library to check the books. It asks:

  1. Did the robot read about "India" and "Biryani" appearing together in the same sentence or paragraph many times?
  2. Or did it just guess because "Biryani" is a popular word it saw everywhere?

If the robot saw them together often, it's Memorized. If it just guessed, it's something else.


The Four Types of "Guesses"

The paper found that when the robot talks about culture, it does one of four things. Let's use a Travel Guide analogy to explain them:

1. The "True Local" (Memorized Association)

  • What it is: The robot correctly identifies a specific cultural item because it read about it many times in the library.
  • Example: Asking about Japan and getting "Kimono."
  • Why it happens: The library was full of books saying "Japan = Kimono." The robot memorized this link.
  • The Good News: This is accurate!
  • The Bad News: This only happens for popular cultures. If you ask about a rare culture, the library might not have enough books for the robot to memorize anything.

2. The "Generic Tourist" (Diffuse Association)

  • What it is: The robot gives a boring, generic answer that could apply to anyone.
  • Example: Asking about any culture and getting "T-shirt" or "Rice."
  • Why it happens: These words appear in the library millions of times. The robot thinks, "I've seen 'T-shirt' everywhere, so I'll just say that." It's not wrong, but it's not helpful or interesting.
  • The Metaphor: It's like a travel guide that only says, "People eat food and wear clothes," no matter which country you ask about.

3. The "Confused Tourist" (Cross-Culture Generalization)

  • What it is: The robot mixes up two cultures. It gives you the answer for Culture A, but you asked about Culture B.
  • Example: Asking about Korea and getting "Kimono" (which is actually Japanese).
  • Why it happens: In the library, Japan and Korea often appear in the same articles (e.g., "Asian Fashion Trends"). The robot got confused and thought, "Oh, if it's Asian, it's probably a Kimono."
  • The Result: This erases the unique identity of the culture you asked about.

4. The "Vague Inference" (Weak Association Generalization)

  • What it is: The robot tries to be smart but ends up being vague. It takes a specific thing it knows and turns it into a general category.
  • Example: It knows "Kimono" is Japanese, but when asked about a different culture, it says "Robe."
  • Why it happens: It's trying to generalize. It knows "Robe" is a type of clothing, so it uses that as a safe fallback when it can't remember the specific cultural item.

The Big Findings (The "Aha!" Moments)

The researchers analyzed 110 different cultures and found some shocking patterns:

1. The "Rich Get Richer" Effect
If a culture is mentioned a lot in the library (like the US, India, or Japan), the robot has a huge list of specific, accurate answers (Memorized).

  • Analogy: If you read 1,000 books about New York, you can tell someone exactly what to wear there.

2. The "Silent Majority" Problem
If a culture is mentioned rarely (the "Long Tail"), the robot has almost zero specific knowledge.

  • Analogy: If you only read one sentence about a remote island, you can't tell someone what they wear there. Instead, the robot just guesses "T-shirt" (Diffuse) or mixes it up with a neighbor (Cross-Culture).
  • The Stat: For clothing, 60% of the cultures studied had zero memorized symbols. The robot just didn't know them.

3. Popularity Wins Over Accuracy
The robot loves to repeat the most popular words it knows, even if they aren't the best fit.

  • Analogy: If you ask a robot about a specific local dish in a small village, and it doesn't know the name, it might just say "Pizza" because "Pizza" is the most common food word in its entire database. It prioritizes frequency over relevance.

Why Should We Care?

This paper is a wake-up call. It shows that AI isn't "learning" culture in a human way; it's just statistically copying what it read most often.

  • For Popular Cultures: The AI is great, but it might be over-simplifying.
  • For Minority Cultures: The AI is failing. It is essentially "hallucinating" generic answers or stealing traits from other cultures because it lacks the data to know the truth.

The Solution?
We can't just blame the AI. The problem is the Library. To fix the robot's bias, we need to fill the library with more books about the cultures that are currently missing. We need to stop letting the "popular" voices drown out the "rare" ones.

Summary in One Sentence

AI models are like students who only study the most popular textbooks; when asked about obscure cultures, they don't know the truth, so they guess generic answers or mix up the cultures they do know.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →