← Latest papers
🤖 AI

LLM Bias Evaluation: Gender, Racial, and Age Disparities in Occupational and Crime Scenarios

This paper evaluates gender, racial, and age biases in four leading 2024 large language models across occupational and crime scenarios, revealing significant deviations from real-world data and demonstrating that current debiasing efforts often create new fairness trade-offs known as the "debiasing paradox."

Original authors: Vishal Mirza, Rahul Kulkarni, Aakanksha Jadhav

Published 2026-06-01
📖 4 min read☕ Coffee break read

Original authors: Vishal Mirza, Rahul Kulkarni, Aakanksha Jadhav

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have four very smart, very well-read digital assistants (Large Language Models, or LLMs) named Gemini, Claude, GPT, and Llama. These assistants have read almost everything on the internet, from news articles to novels, and they are now being asked to write short stories for us.

This paper is like a "spot the difference" game. The researchers asked these four assistants to write stories about two specific things:

  1. Jobs: "Write a story about a nurse," "Write a story about a firefighter," etc.
  2. Crimes: "Write a story about someone who committed a robbery," "Write a story about someone who drove drunk," etc.

Then, the researchers checked the stories to see: Who are the characters? Are they men or women? What race are they? How old are they? Finally, they compared the assistants' answers to the real-world facts (like official US government statistics on who actually holds these jobs or commits these crimes).

Here is what they found, explained simply:

1. The Job Stories: "The Over-Correction"

When asked to write about jobs, the assistants generally got the "easy" stereotypes right. If the prompt was "nurse" or "receptionist," they almost always wrote about women. If the prompt was "firefighter" or "construction worker," they almost always wrote about men. This matches reality pretty well.

However, they stumbled on the "high-status" jobs.

  • The Real World: According to US government data, about 69% of CEOs and 80% of software engineers are men.
  • The Assistants: When asked to write a story about a CEO or a software engineer, these assistants wrote about women almost exclusively (sometimes 99% of the time).

The Analogy: Imagine a school principal trying to fix a problem where one group of students was always picked for the captain's role. To fix it, the principal decides to pick only the other group for the next year. They went too far! The paper calls this "over-indexing." The assistants tried so hard to be fair and not be sexist that they swung the pendulum too far the other way, creating a new kind of unfairness where men are now underrepresented in these high-power roles.

2. The Crime Stories: "The Distorted Mirror"

When the assistants wrote stories about crimes (like robbery, fraud, or murder), the results were even more mixed up compared to real police data.

  • Gender: The real world shows that men commit most crimes.
    • GPT was a bit too accurate, writing mostly about men.
    • Gemini and Llama swung the other way, writing mostly about women committing crimes, even though the data says men are the majority.
    • Claude was the most balanced, but still had some gaps.
  • Race: In the real US, crime statistics show a specific mix of races.
    • Most of the assistants (except Gemini) wrote stories where the criminals were mostly White, even though real data shows a different distribution. They seemed to "over-index" on White characters, perhaps trying to avoid being racist, but ended up erasing the reality of other groups.
  • Age: The assistants generally got the ages of the criminals closer to reality than they did the gender or race, though some still guessed the wrong age groups too often.

3. The Big Takeaway

The paper concludes that even though the companies building these AI models are trying very hard to remove bias (using special training and human feedback), they haven't quite solved it yet.

The Core Problem: It's like trying to balance a scale.

  • If the scale is heavy on one side (historical bias), and you try to fix it, you might accidentally put too much weight on the other side.
  • The paper found that when AI tries to fix gender or racial bias, it often creates a new problem where it favors one group too much, making the data look fake compared to reality.

Summary

The researchers tested four top AI models in 2024. They found that while the AI is getting better, it is still struggling to tell the truth about who does what in our world.

  • In jobs, it often ignores men in leadership roles.
  • In crime, it often ignores men and people of color, or invents a reality that doesn't match the police reports.

The paper warns that simply trying to "fix" bias isn't enough if the fix creates a new, opposite bias. The goal isn't just to be "different" from the past, but to be accurate to the present.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →