← Latest papers
🤖 machine learning

It's the humans, not the data: Geopolitical bias in LLMs originates in post-training, amplified by the language of the prompt

This study demonstrates that geopolitical bias in large language models is primarily introduced and amplified during the post-training alignment phase rather than inherited from pre-training data, with the direction and magnitude of these biases varying by the model developer's region and the language used in prompting.

Original authors: Stuart Bladon, Brinnae Bent

Published 2026-05-25
📖 5 min read🧠 Deep dive

Original authors: Stuart Bladon, Brinnae Bent

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have a group of students from different countries (the USA, France, and China) who are all learning to write essays about world events.

For a long time, people assumed that if a student wrote a biased essay, it was because they had read biased books in the library. In the world of AI, this "library" is the massive amount of internet data the model reads before it starts learning to chat. This is called pre-training.

However, this paper argues that the real story is different. It's not what the students read in the library; it's what happens in the final exam prep class (called post-training) that changes their minds.

Here is the breakdown of their findings using simple analogies:

1. The "Blank Slate" vs. The "Final Exam"

The researchers looked at seven different AI "families" (like different schools). For each family, they compared two versions:

  • The Base Model: This is the student right after reading the library books but before any specific coaching.
  • The Chat Model: This is the same student after going through a "coaching" phase (post-training) where humans teach them how to answer questions politely and safely.

The Finding: The "Base" students were mostly neutral. They didn't really have a strong opinion on which country was right or wrong in a conflict. But once they went through the "coaching" (post-training), they suddenly developed strong opinions.

  • The Analogy: Imagine a student who is neutral about a sports rivalry. After their coach tells them, "You must support Team A," the student suddenly starts cheering loudly for Team A. The bias didn't come from the books they read; it came from the coach's instructions.

2. The "Coach's Nationality" Matters

The researchers found that the direction of the bias depended on where the coach was from.

  • If the AI was developed by a Western lab (USA or France), the coaching made the AI lean against China in geopolitical scenarios.
  • If the AI was developed by a Chinese lab (like Alibaba), the coaching made the AI lean in favor of China.
  • The Exception: One Chinese lab (01.AI) actually coached their student to lean against China, proving that it's not just about the country, but the specific choices made by the people building the model.

The Big Shift: The most dramatic example was Alibaba's Qwen model.

  • Before coaching: It was neutral (like a blank page).
  • After coaching: It became extremely pro-China. The researchers say the odds of it favoring China jumped 18 times higher just because of the post-training.

3. The "Language of the Question" Amplifies the Bias

The researchers also tested if the language used to ask the question changed the answer.

  • The French Example: The French AI (Mistral) was neutral when asked in English. But when asked in French, it suddenly became very pro-France.
  • The Chinese Example: Asking Chinese-made models in Chinese sometimes made them more pro-China, but not always. Sometimes the language acted like a volume knob, turning the bias up or down depending on the specific model.

The Analogy: Think of the AI as an actor. If you ask the actor a question in English, they might play a neutral character. But if you whisper the same question in their native language, they might suddenly start acting out a role that fits their cultural background much more strongly.

4. It's Not Just About "Real" Countries

To make sure the AI wasn't just memorizing facts (like "China is big"), the researchers invented fake countries with fake names (like "Zhaodong" which sounds Chinese, or "Bretherland" which sounds English).

  • The Result: Even with fake countries, the Chinese AI favored the one that sounded Chinese, and the French AI favored the one that sounded French.
  • The Lesson: The AI isn't just recalling facts; it's applying a cultural lens to any situation, real or fake, based on how it was trained to think.

5. The "Refusal" Trick

Some Chinese models initially seemed to refuse to answer questions about China (saying "I can't answer that"). The researchers realized this wasn't a lack of opinion; it was a formatting trick. When they looked deeper, they found that once the model was allowed to write a full sentence, it did have a strong opinion.

  • The Analogy: It's like a student who raises their hand and says, "I can't answer," but when the teacher says, "Just write down one word," they immediately write "Team A." The opinion was there all along; the format just hid it.

The Bottom Line

The paper concludes that geopolitical bias in AI is not an accident of the data it was fed. Instead, it is actively built during the final training stage where humans align the AI with specific values.

  • Old Belief: "The AI is biased because the internet is biased."
  • New Finding: "The AI is biased because the people who trained it (the coaches) shaped it to have those specific preferences."

This means that if we want to fix bias in AI, we can't just clean up the internet data. We need to look at the alignment process—the specific rules and human feedback used to teach the AI how to behave.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →