← Latest papers
💬 NLP

Large Language Models Reproduce Racial Stereotypes When Used for Text Annotation

This paper demonstrates that across 19 large language models and over 4 million annotation judgments, subtle identity cues such as names and dialects systematically trigger racial stereotypes, causing texts associated with minority groups to be rated as more aggressive, less professional, and less confident, thereby embedding these biases into automated datasets used for research and decision-making.

Original authors: Petter Törnberg

Published 2026-03-17
📖 5 min read🧠 Deep dive

Original authors: Petter Törnberg

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you hire a team of 19 different robots to help you read thousands of short stories about people. Your goal is to have these robots act as "judges," rating the characters in the stories on things like: Are they smart? Are they angry? Would you hire them for a job? Are they professional?

You tell the robots, "Just look at the words and give us a score." You don't tell them who the people are.

But here's the twist: The robots have read the entire internet to learn how to speak. Unfortunately, the internet is full of old, unfair stereotypes about race. So, even though you didn't tell the robots the race of the characters, the robots figured it out on their own just by looking at tiny clues—like a person's name or the way they spoke.

This paper is a massive investigation into what happens when these robots judge people based on those clues. The researchers found that the robots didn't just make random mistakes; they reproduced the exact same racial stereotypes that humans have held for decades.

Here is the breakdown of what happened, using some simple analogies:

1. The "Name Game" (Experiment 1)

The researchers took the exact same story and just swapped the name at the top.

  • The Black Name Effect: When the story had a name like "Tyrone" or "Shaniqua," the robots almost always rated the person as more aggressive and more likely to gossip. It's as if the robot saw a name and immediately put a "troublemaker" sticker on the file, even though the story was identical to one with a white name.
  • The "Bamboo Ceiling" for Asian Names: When the name was Asian (like "Wei" or "Mei"), the robots got confused in a specific way. They rated these people as super smart and creative (the "good student" stereotype). But, they also rated them as less confident, less sociable, and less likely to be a leader. It's like the robots thought, "Great at math, but don't put them in charge."
  • The Arab Name Effect: These names triggered a mix. The robots thought these people were smart and ambitious, but also less warm, more emotional, and less good at communicating.
  • The "Self-Discipline" Penalty: No matter which minority name was used (Black, Asian, Arab, or Hispanic), the robots almost universally rated them as less self-disciplined than white names. It was a consistent "penalty" applied to everyone.

The One Weird Exception:
When it came to hiring, the robots did the opposite of what humans usually do in real life. In real life, studies show that resumes with minority names get fewer callbacks. But these robots? They actually favored the minority names for hiring!

  • Why? The researchers think the companies that built these robots (like OpenAI and Google) specifically trained them to be "nice" and avoid hiring discrimination. They over-corrected. They fixed the "hiring" bias so hard that they swung the pendulum the other way, but they didn't fix the deeper, invisible biases about personality or character.

2. The "Accent" Test (Experiment 2)

Next, the researchers tried a different trick. They kept the name the same (a white name) but changed the dialect of the text.

  • Version A: Written in standard English (like a news report).
  • Version B: Written in African American Vernacular English (AAVE), using different grammar and slang, but saying the exact same thing.

The Result: This was the most shocking part. Every single one of the 19 robots judged the AAVE version much more harshly.

  • They thought the AAVE speaker was less professional.
  • They thought they were less educated.
  • They thought they were angrier and more toxic.
  • They were less likely to be hired.

The Analogy: Imagine two people walking into a room wearing the exact same suit. One speaks with a standard accent, and the other speaks with a regional dialect. If you ask a robot to judge them, the robot will say the person with the dialect is "unprofessional" and "angry," even though they are wearing the same suit and saying the same words. The robots are punishing the voice, not the content.

Why Does This Matter?

You might think, "Okay, robots are biased. So what? We just won't use them."

The problem is that we are using them.

  • Universities use them to grade essays.
  • Companies use them to screen resumes.
  • Social media companies use them to decide what posts are "toxic" and should be deleted.

If a robot is used to moderate content, it might delete a post written in AAVE because it thinks the post is "angry" or "toxic," while letting a post with the same meaning written in standard English slide. If a robot is used to screen job applicants, it might unfairly reject candidates based on subtle clues in their writing style.

The Big Takeaway

These Large Language Models are like mirrors. They reflect the world they were trained on. Because our world has deep-seated racial biases, these mirrors show us those biases back to us, often in ways we don't expect.

The researchers warn us: We cannot just trust these robots to be neutral. Even if we try to "fix" them, they might fix one problem (like hiring) while leaving others (like judging personality or dialect) broken. If we use these tools to make important decisions about people's lives, we risk automating and speeding up old-fashioned racism, making it look like "neutral math."

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →