← Latest papers
💬 NLP

Probing Social Identity Bias in Chinese LLMs with Gendered Pronouns and Social Groups

This study evaluates social identity biases in ten Chinese large language models using Mandarin-specific prompts to compare ingroup and outgroup framings across 240 social groups, revealing that while instruction tuning reduces sentiment asymmetries, toxicity gaps persist and are further exacerbated by the use of explicitly feminine plural pronouns.

Original authors: Geng Liu, Feng Li, Junjie Mu, Mengxiao Zhu, Francesco Pierri

Published 2026-05-28
📖 4 min read☕ Coffee break read

Original authors: Geng Liu, Feng Li, Junjie Mu, Mengxiao Zhu, Francesco Pierri

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have a group of very smart, digital storytellers (Large Language Models, or LLMs) that speak Chinese. These models have read almost everything written on the internet, so they know a lot about the world. But, just like people, they might have picked up some hidden prejudices along the way.

This paper is like a detective story where researchers put these digital storytellers to a test to see if they treat "us" differently than "them."

The Setup: "Us" vs. "Them"

The researchers wanted to see if the models would be nicer to a group when the story was told from the inside ("We are...") compared to when it was told from the outside ("They are...").

Think of it like a school playground. If a student says, "We are the soccer team," they might sound proud and happy. But if they say, "They are the soccer team" (referring to the rival team), they might sound critical or mean. The researchers asked: Do these AI models act like that?

They tested 10 different Chinese AI models using 240 different social groups (like people of different ages, jobs, or backgrounds).

The Special Chinese Twist: The "He" vs. "She" Pronoun

Here is where the study gets really clever. In English, the word "they" is used for a group of people regardless of gender. But in Chinese, there is a visual difference in the writing:

  • 他们 (Tāmen): The default "they" for a mixed group or when you don't know the gender. It looks like a "person" with a "horse" (a neutral symbol).
  • 她们 (Tāmen): The specific "they" for an all-female group. It looks like a "person" with a "female" symbol.

The researchers used this difference as a special tool. They asked the AI to talk about groups using the neutral "they" and then the female-specific "they." They wanted to see if the AI got meaner just because the group was explicitly marked as female.

What They Found

1. The "Us" is Better Than "Them" Bias
Just like the playground analogy, the AI models generally liked the "We" groups more than the "They" groups.

  • When the AI said "We are...", the responses were warmer and more positive.
  • When the AI said "They are...", the responses were colder and sometimes even hostile.
  • The Catch: This wasn't a huge explosion of hate, but a subtle, consistent pattern. It's like a slight frown when talking about outsiders compared to a smile for insiders.

2. The "Teacher" vs. The "Student" (Instruction Tuning)
The researchers tested two types of models:

  • Base Models: These are like raw students who have read a lot but haven't been taught how to behave politely yet. These models were much more likely to be mean to the "They" groups.
  • Instruction-Tuned Models: These are like students who have been taught by teachers (humans) to be helpful and safe. These models were better at being polite to everyone. They reduced the "smile vs. frown" gap.
  • However: Even the "good students" (the polite ones) still had a hidden problem. While they stopped being mean in their tone, they still showed signs of toxicity (rude or unsafe language) when talking about "They" groups, especially if the group was marked as female.

3. The "Female" Penalty
This was a surprising discovery. In several models, when the AI used the female-specific pronoun (她们), the responses were actually more toxic (ruder or more harmful) than when it used the neutral pronoun (他们).

  • It's as if the AI, even when trying to be polite, subconsciously thought, "Oh, this group is all women," and immediately became a little more critical or unsafe in its language. This is a bias that you wouldn't see in English because English doesn't have a visual difference for "they."

4. Real Life Check
The researchers also looked at real conversations between humans and AI (from a public dataset). They found that even in real life, the AI tended to be nicer to "us" and slightly more critical of "them," suggesting this isn't just a lab experiment—it happens in the real world too.

The Bottom Line

The paper concludes that Chinese AI models aren't neutral robots. They carry social biases:

  • They prefer "Us" over "Them."
  • They can be meaner to "Them" than they are nice to "Us."
  • They have a specific, hidden bias against groups marked as female, showing up as rude language even when the AI tries to be polite.

The researchers say that to fix this, we can't just use English rules. We need to understand the specific quirks of the Chinese language (like those pronoun differences) to make these AI tools fair and safe for everyone.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →