← Latest papers
💬 NLP

The Digital Afterlife of Empires: Four Language Models Converge on the Same Imperial Cartography of Writing

This paper argues that historical imperial inequalities persist in contemporary large language models through shared training data, resulting in radical digital under-support for most writing systems and systematic, convergent errors across diverse model architectures that disproportionately misattribute non-religious scripts to religious use.

Original authors: Hiroki Fukui

Published 2026-06-30
📖 5 min read🧠 Deep dive

Original authors: Hiroki Fukui

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

The Big Idea: The Ghost in the Machine

Imagine that history is a giant library. For 500 years, powerful empires (like Spain, Britain, or France) decided which books to keep, which to burn, and which languages to write in. They destroyed many local writing systems and forced others to change.

This paper argues that Large Language Models (LLMs)—the AI brains behind tools like ChatGPT—are not new, neutral inventions. Instead, they are like ghosts haunting a library built by those empires. Even though the empires are gone, their "ghosts" are still in the building. The AI doesn't know how to read certain languages not because the engineers were mean, but because the library (the internet data) they learned from is missing those books.

The Main Findings, Broken Down

1. The "Digital Tax" (Why some languages cost more)

Imagine you go to a toll booth to enter a highway.

  • English speakers pay $1 to drive 10 miles.
  • Speakers of minority languages (like Limbu or Adlam) are charged $31.70 to drive the exact same 10 miles.

Why? The AI breaks text down into small chunks called "tokens" (like Lego bricks).

  • For English, the AI has a huge box of pre-made Lego bricks. It snaps them together easily.
  • For minority scripts, the AI doesn't have the right bricks. It has to break the letters down into raw, meaningless bytes (like smashing the Lego bricks into plastic dust and trying to rebuild them from scratch).
  • The Result: It takes 31 times more "computing effort" to process a sentence in Limbu than in English. This isn't a bug; it's a "tax" built into the system because those languages were historically suppressed and have less data on the internet.

2. The "Digital Funnel" (Who gets to be heard)

The researchers looked at 300 different writing systems in human history. They tried to see how well the digital world supports them.

  • The Funnel: Imagine a giant funnel.
    • Top: 300 scripts enter.
    • Middle: Many get stuck because they aren't in the computer's dictionary (Unicode) or can't be read by scanners (OCR).
    • Bottom: Only 29 scripts (less than 10%) make it all the way through to the bottom.
  • The Pattern: The scripts that make it through are almost exclusively the ones used by historical empires (Latin, Chinese, Arabic, Cyrillic). The scripts that get stuck are the ones those empires tried to erase.
  • The Tragedy: 60 languages are still spoken by living people today, but the digital world treats them as if they are dead. They are "living but digitally dead."

3. The "Four Friends" Experiment (It's not just one AI)

To prove this wasn't just a mistake by one company (like OpenAI), the researchers asked four different AI families (Claude, GPT-4o, Grok, and DeepSeek) the same 3,000 questions about writing systems.

  • The Surprise: These four AIs are built by different companies, in different countries, with different code. They shouldn't agree on mistakes.
  • The Reality: They agreed on their mistakes 90% of the time.
  • The Analogy: Imagine four students who went to four different schools. If they all get the same wrong answer on a history test, it's not because they copied each other. It's because they all read the same biased textbook.
  • The Bias: The AIs consistently guessed wrong in the same way. For example, if a script was from a non-Western country, the AIs were very likely to guess, "Oh, this must be used for religion," even if it wasn't. They were filling in the blanks with stereotypes found in the data.

4. The "Empire Without an Emperor"

The most important finding is that nobody is currently in charge of this bias.

  • The engineers didn't sit down and say, "Let's make Limbu speakers pay more."
  • The bias is structural. It happened like this:
    1. Empires destroyed communities and reduced the number of speakers.
    2. Fewer speakers meant fewer books and websites written in those languages.
    3. Fewer websites meant the AI had no data to learn from.
    4. No data means the AI is bad at that language.
  • The "Empire" doesn't need a leader anymore. The damage is baked into the statistics. The AI is just doing math based on a world that was already unequal.

The "Religion" Glitch

One specific finding stood out: The AIs were obsessed with guessing that scripts were "religious."

  • Out of all the mistakes the four AIs made together, 43% were about guessing a script was used for religion when it wasn't.
  • Why? The internet is full of old, colonial-era descriptions of non-Western cultures that focused heavily on their "mystical" or "religious" aspects, while ignoring their daily, administrative, or scientific uses. The AI learned that "Non-Western Script = Religion."

The Conclusion: What Can Be Done?

The paper suggests that we can't just "fix" the AI by tweaking its code.

  • The Problem: The problem is the library (the training data), not the reader (the AI).
  • The Solution: To fix this, we need to change the library. We need to add more books written by the people who speak these languages, in their own scripts, describing their own lives.
  • The Warning: Until we do that, the AI will continue to be a mirror of the last 500 years of history, reflecting the same inequalities, even if no one intended for it to happen.

In short: The AI isn't broken; it's just faithfully repeating the history of who was allowed to write and who was silenced. The "digital afterlife" of empires is still very much alive.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →