← Latest papers
💬 NLP

LLMs as Cultural Archives: Cultural Commonsense Knowledge Graph Extraction

This paper introduces an iterative framework for extracting Cultural Commonsense Knowledge Graphs from large language models to structure implicit cultural knowledge, revealing that while these graphs enhance cultural reasoning and generation, they currently exhibit significant bias toward English even when representing non-English cultures.

Original authors: Junior Cedric Tonga, Chen Cecilia Liu, Iryna Gurevych, Fajri Koto

Published 2026-01-27
📖 5 min read🧠 Deep dive

Original authors: Junior Cedric Tonga, Chen Cecilia Liu, Iryna Gurevych, Fajri Koto

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine that Large Language Models (LLMs) are like massive, invisible libraries. These libraries are filled with books written by people all over the world, so they contain a huge amount of information about how different cultures think, act, and celebrate. However, right now, this information is messy. It's buried deep inside the library, scattered in millions of unorganized sentences, making it hard to find or understand the specific "rules" of a culture.

This paper proposes a way to clean up that library and turn it into a Cultural Commonsense Knowledge Graph (CCKG). Think of this graph not as a list of facts, but as a train track system.

The Core Idea: Turning Facts into Train Tracks

Usually, if you ask an AI about culture, it might give you a single fact, like "Indonesians eat soto ayam for breakfast." That's a static fact.

The authors wanted to map out the journey of that fact. They built a system that asks the AI: "If someone wants breakfast, what do they do next? Then what? And what happens after that?"

  • The Track: "If you look for breakfast \rightarrow you go to a warung (small shop) \rightarrow you order soto ayam \rightarrow you feel warm and comforted."

By connecting these steps, they create a "train track" of cultural logic. This allows the AI to understand not just what happens, but the flow of cultural life.

How They Built the Tracks (The Process)

The researchers treated the AI (specifically GPT-4o) as a "cultural archivist." They used a two-step process to build these tracks:

  1. Laying the First Rails (Initial Generation): They asked the AI to generate simple "If-Then" statements about daily life topics (like food, weddings, or habits) for five different countries: China, Indonesia, Japan, England, and Egypt.
  2. Extending the Tracks (Iterative Expansion): This is the clever part. Once the AI said, "If you look for breakfast, you order soto ayam," the researchers asked the AI to fill in the gaps.
    • Intermediate Expansion: "What happens between looking for breakfast and ordering? Maybe you choose a side dish first."
    • Forward Expansion: "What happens after ordering? Maybe you add sambal (chili sauce)."

They repeated this process, turning a single fact into a long, logical chain of cultural behavior.

The Big Surprise: The "English Bridge"

The researchers had a hypothesis: "If we want to know about Indonesian culture, we should ask the AI in Indonesian. If we want to know about Chinese culture, we should ask in Chinese."

They were wrong.

When they tested the quality of these cultural tracks, they found a strange but consistent pattern: The tracks built in English were much better than the tracks built in the native languages.

  • The Analogy: Imagine trying to draw a map of a city. You might think drawing it in the local language would be most accurate. But the researchers found that the AI's "internal map" was drawn most clearly in English. Even when describing Indonesian breakfasts or Egyptian weddings, the AI produced more logical, coherent, and culturally accurate "tracks" when it was forced to speak English.
  • The Result: The English versions of the knowledge graph were more consistent and made more sense to human evaluators than the versions written in Chinese, Arabic, or Indonesian.

Does This Help Smaller AIs?

The researchers then took these "train tracks" (the CCKG) and gave them to smaller, weaker AI models to see if it helped them understand culture better.

  • The Test: They asked these smaller models to answer questions about culture or write short stories.
  • The Outcome: When the smaller models were allowed to "look at the tracks" (using the CCKG as a guide), they got much better at answering questions and writing stories that felt culturally authentic.
  • The Catch: The biggest boost came when the models used the English tracks, even for non-English cultures. It seems that the "English bridge" is currently the most reliable way to access the AI's cultural knowledge.

What They Did NOT Claim

It is important to stick to what the paper actually says:

  • They did not say this system is ready for use in hospitals, schools, or legal systems.
  • They did not claim they solved the problem of AI bias or stereotypes (in fact, they noted that bias is still a risk and they didn't deeply audit for it).
  • They did not say that English is the "best" language for humans to speak; they only found that, for these specific AI models, English is the most stable language for extracting cultural logic.

Summary

The paper shows that we can treat AI models as giant archives of cultural knowledge. By asking the AI to build "train tracks" of cultural logic (If this happens, then that happens), we can create a structured map of how different cultures work. However, there is a twist: currently, these AI models seem to understand and express these cultural maps most clearly when they are speaking English, even when the culture itself is not English-speaking. This "English advantage" helps smaller AI models perform better, but it also highlights that our current AI tools are still uneven in how they represent the world's diverse cultures.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →