← Latest papers
🤖 AI

Generating the Modal Worker: A Cross-Model Audit of Race and Gender in LLM-Generated Personas Across 41 Occupations

This paper audits four major large language models across 41 U.S. occupations and finds that they systematically distort racial and gender demographics by compressing occupational profiles toward dominant stereotypes, significantly underrepresenting White and Black workers while overrepresenting Hispanic and Asian workers, thereby revealing shared structural biases that reshape demographic visibility in synthetic populations.

Original authors: Ilona van der Linden, Sahana Kumar, Arnav Dixit, Aadi Sudan, Smruthi Danda, David C. Anastasiu, Kai Lukoff

Published 2026-03-30
📖 5 min read🧠 Deep dive

Original authors: Ilona van der Linden, Sahana Kumar, Arnav Dixit, Aadi Sudan, Smruthi Danda, David C. Anastasiu, Kai Lukoff

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have four different "Digital Storytellers" (AI models like GPT-4, Gemini, DeepSeek, and Mistral). You ask them a simple question: "Tell me about a person who works as a [Job Title] in the United States."

You ask this for 41 different jobs, from nurses to engineers to truck drivers, and you ask each storyteller to create 10,000 different people for each job. That's over 1.5 million made-up people.

The researchers in this paper acted like detectives. They took these 1.5 million made-up characters and compared them to the real-world census data (the U.S. Bureau of Labor Statistics) to see if the AI was telling the truth about who actually does these jobs.

Here is what they found, explained simply:

1. The "One-Size-Fits-All" Problem (The Mold)

When you ask a human to describe a "nurse," they might say, "It could be a woman, a man, a young person, or an older person."

But these AI models act like cookie cutters.

  • The Finding: For most jobs, the AI doesn't create a mix of people. It picks one dominant type and stamps it out 99% of the time.
  • The Analogy: Imagine a bakery that makes "Engineer" cookies. Instead of making a variety of flavors, the bakery decides every single engineer cookie must be chocolate. If you ask for 1,000 engineer cookies, you get 1,000 chocolate ones. The AI does this with gender and race. It stops being a "population" and starts being a "stereotype."

2. The "Stereotype Amplifier" (The Funhouse Mirror)

The AI doesn't just copy reality; it often exaggerates it, like a funhouse mirror that stretches things out.

  • The Finding: If a job is already mostly men in real life (like a truck driver), the AI makes it look even more like a man's job. If a job is mostly women (like a preschool teacher), the AI makes it look almost exclusively women.
  • The Result: The "middle ground" disappears. The AI creates a world where jobs are extremely segregated, making it harder for people to imagine themselves in roles that don't fit the "cookie cutter."

3. The "Invisible People" and the "Over-Represented"

The researchers found some very specific, weird patterns in how the AI handled race:

  • The "Ghosting" of Black Workers: In many jobs, the AI almost completely erased Black workers. If you asked for a "Librarian" or a "Biologist," the AI rarely imagined a Black person, even though they exist in those jobs in real life. It's like the AI decided those people simply don't exist in those rooms.
  • The "Hispanic Housekeeper" Trope: The AI was obsessed with making housekeepers Hispanic. In real life, housekeepers are diverse. In the AI's world, they were nearly 100% Hispanic. It took a real trend and turned it into a caricature.
  • The "Asian Professional" Boost: Conversely, the AI loved putting Asian faces in high-status jobs like scientists, librarians, and doctors, often making them appear more frequently than they actually do in the real workforce.

4. The "Universal Bias" (It's Not Just One Bad Apple)

You might think, "Well, maybe the American AI (GPT-4) is biased, but the Chinese one (DeepSeek) or the French one (Mistral) is different."

  • The Shock: No. All four models, from different countries and companies, made the same mistakes.
  • The Analogy: It's like if you asked four different chefs from four different countries to make a "Classic Burger," and they all accidentally put pickles on the ice cream instead of the burger. This suggests the problem isn't just one chef's bad recipe; it's that they all learned from the same library of books (the internet) which already contains these stereotypes.

5. The "Salary Glitch"

The AI also tried to guess salaries.

  • The Finding: The AI knew that CEOs make more than janitors. But when it came to men vs. women, the AI mostly ignored the real-world pay gap. It gave men and women the exact same salary in the same job.
  • The Twist: While this sounds "fair," the researchers argue it's actually a problem. By pretending the pay gap doesn't exist, the AI hides a very real, very unfair problem that exists in the real world.

Why Does This Matter?

Think of these AI personas as digital role models.

  • If a young girl asks an AI, "What does a software engineer look like?" and the AI shows her a white man 99% of the time, she might subconsciously think, "That job isn't for me."
  • If a young Black boy asks, "What does a librarian look like?" and the AI shows him a white woman, he might feel invisible.

The Bottom Line:
These AI tools are currently acting like lazy, biased photographers. Instead of taking a wide, diverse photo of the workforce, they are taking a narrow, distorted selfie that reinforces old stereotypes. The paper argues that we need to fix the "camera settings" so that when AI generates people, it reflects the messy, diverse, and complex reality of the real world, not a simplified cartoon version of it.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →