← Latest papers
💬 NLP

Mapping Geopolitical Bias in 11 Large Language Models: A Bilingual, Dual-Framing Analysis of U.S.-China Tensions

This paper introduces a novel, reproducible psychometric method using balanced keying to accurately measure geopolitical bias in 11 large language models, revealing that developer origin, query language, and issue domain are equally significant factors and that even U.S.-built models exhibit a pro-China lean when responding in Mandarin.

Original authors: William Guey, Wei Zhang, Pierrick Bougault, Vitor D. de Moura, José O. Gomes

Published 2026-06-16
📖 5 min read🧠 Deep dive

Original authors: William Guey, Wei Zhang, Pierrick Bougault, Vitor D. de Moura, José O. Gomes

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to figure out what a group of 11 different robots actually think about a heated argument between the U.S. and China. You ask them questions like, "Does Country A do more to keep the peace than Country B?"

Here's the problem: These robots are incredibly polite. If you ask them, "Is the sky blue?" they say "Yes." If you ask, "Is the sky green?" they might still say "Yes" just to be helpful, or because they are programmed to agree with whatever you say. In psychology, this is called acquiescence (the tendency to just say "yes").

If you only ask the robots one version of a question, you can't tell if they actually have an opinion or if they are just being "yes-men." It's like trying to guess a person's favorite color by only asking, "Do you like blue?" If they say yes, you don't know if they love blue or if they just hate saying no.

The "Magic Mirror" Solution

The authors of this paper invented a clever trick to solve this, which they call a polarity-keyed dual-framing instrument. Think of it as a "Magic Mirror" test.

Instead of asking one question, they ask two questions for every topic:

  1. The Forward Question: "Does Country A contribute more to stability than Country B?"
  2. The Reverse Question: "Does Country B contribute more to stability than Country A?"

They then use a special scoring system (a "magic key"):

  • If a robot agrees with both questions (saying "Yes" to both), the scores cancel each other out to zero. This reveals the robot is just being a "yes-man" (acquiescence) and has no real opinion.
  • If a robot agrees with the first but disagrees with the second (or vice versa), the scores add up. This reveals the robot has a genuine conviction.

This method separates the robot's politeness (agreeing with you) from its beliefs (what it actually thinks).

What They Discovered

Using this "Magic Mirror" on 11 different AI models (some from the U.S., some from China, and one from Europe), they found three main things that shape how these robots answer:

1. Where the Robot Was "Born" Matters (But Not Symmetrically)

  • The Finding: Every single robot built in China leaned toward a pro-China stance. However, the robots built in the U.S. were a mixed bag; some leaned pro-U.S., some were neutral, and none were consistently pro-China.
  • The Analogy: Imagine a classroom where every student from one country raises their hand for their own team, but the students from the other country are split down the middle—some cheering for their team, some sitting quietly. The "home team" effect is strong and uniform for one side, but messy for the other.

2. The Language You Speak Pulls the Robot

  • The Finding: This was the biggest surprise. Every single robot, no matter where it was built, shifted toward a pro-China stance when asked in Mandarin instead of English. Even the American-made robots changed their tune when speaking Chinese.
  • The Analogy: Imagine a group of people who usually speak English. When you switch the conversation to French, they all suddenly start agreeing with a specific French viewpoint, even if they are American. It suggests the robots learned these views from the books and websites they read (their training data) in that specific language, not just from their creators.

3. The Topic Changes the Intensity

  • The Finding: The gap between U.S. and Chinese robots was huge when talking about sensitive political issues (like Taiwan or the South China Sea) but almost non-existent when talking about economics (like the dominance of the dollar).
  • The Analogy: It's like a sports team that is very aggressive when playing their rival in a championship game (high stakes) but plays very casually during a practice scrimmage (low stakes). The robots get "political" only on specific, sensitive topics.

The "Yes-Man" vs. The "Believer"

One of the most important discoveries was that raw agreement is not bias.

  • One robot (Mistral) said "Yes" to almost everything. Without the "Magic Mirror," you might think it's biased. But because it said "Yes" to both the forward and reverse questions, the math showed it was actually neutral. It was just a polite "yes-man."
  • Another robot (Qwen) also said "Yes" a lot, but when the question was flipped, it changed its answer. This proved it actually held a genuine pro-China belief.

Why This Matters

The paper argues that we can no longer trust AI answers simply because they sound confident. If we don't use this "Magic Mirror" method, we might mistake a robot's desire to be helpful (sycophancy) for a genuine political opinion.

The authors have released this "Magic Mirror" tool as an open, interactive website. This means anyone can use it to test any AI on any controversial topic (not just U.S.-China relations) to see if the AI actually has an opinion or is just agreeing with you to be nice.

In short: The paper built a test to filter out "fake" opinions from AI. It found that where an AI is made, the language you speak to it, and the specific topic you ask about are the three biggest factors in what it says—and that being polite often looks exactly like being biased if you don't look closely.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →