← Latest papers
💬 NLP

Can LLMs Hire Fairly? Racial Bias in Resume Screening

This paper audits fourteen large language models and finds that while a 2023 model exhibited pro-White hiring bias similar to human discrimination, all models released in 2024 or later demonstrated either no bias or a significant pro-Black reversal, marking a generational shift in algorithmic hiring fairness.

Original authors: Zhenyu Gao, Wenxi Jiang, Yutong Yan

Published 2026-06-30
📖 4 min read☕ Coffee break read

Original authors: Zhenyu Gao, Wenxi Jiang, Yutong Yan

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are hiring for a job, but instead of a human looking at the resumes, you ask a super-smart computer robot to make the decision. You want to know: Is this robot fair?

This paper is like a massive "mystery shopper" experiment. The researchers hired 14 different AI robots (Large Language Models) to act as hiring managers. They gave each robot thousands of pairs of fake resumes that were identical in every way—same education, same work history, same age, same skills. The only thing that changed was the name on the resume.

Here is the story of what they found, told through a few simple analogies:

1. The "Old Robot" vs. The "New Robots"

Think of the AI models like different generations of smartphones.

  • The 2023 Model (The "Old Phone"): The researchers tested one older model (GPT-3.5-turbo). This robot acted exactly like the biased humans found in past studies. When it saw a name that sounded Black, it was less likely to say "Yes, call this person in" compared to a name that sounded White. It showed a clear preference for White names.
  • The 2024+ Models (The "New Phones"): Then, they tested 13 newer models released in 2024 and 2025. These robots did something surprising. They didn't just become neutral; they swung the other way. They started favoring Black names over White names.

The Analogy: Imagine a scale.

  • The old robot tipped the scale heavily to the White side.
  • The new robots tipped the scale heavily to the Black side.
  • The newest robots (from 2026) are trying to balance the scale perfectly, showing almost no preference either way.

2. The Same Story for Men and Women

They ran the exact same test but changed the names to see if the robots were biased against men or women.

  • The Old Robot: Preferred men over women.
  • The New Robots: Preferred women over men.
  • The Newest Robots: Are becoming neutral again.

3. Why Did This Happen?

The paper suggests this isn't a glitch; it's a result of how the robots were "taught" to behave.

  • The Training Data: The old robot learned from the internet as it was, which contains a lot of historical human bias (favoring White men).
  • The "Alignment" Fix: The companies that built the newer robots realized this was unfair. They added a special "training class" (called alignment) to teach the robots to be fair.
  • The Over-Correction: It seems the newer robots got a little too excited about being fair. Instead of just being neutral, they started actively favoring the groups that were previously discriminated against. It's like a teacher who, after seeing a student get bullied, decides to give that student extra credit on every test, even when they didn't earn it.

4. The Big Takeaway

The main lesson from this paper is that AI bias is not a fixed thing. It changes depending on which version of the robot you use and when it was built.

  • 2023 Robots: Repeated the old human mistakes (favoring White men).
  • 2024-2025 Robots: Tried to fix it but swung too far the other way (favoring Black women).
  • 2026 Robots: Are getting closer to being truly neutral.

The Bottom Line:
The paper concludes that just because a tool is "AI" doesn't mean it's fair. Sometimes it's too biased one way, and sometimes it's too biased the other. The only way to know if a specific AI hiring tool is fair is to test it constantly, because the "personality" of the AI changes with every new update.

What the paper does NOT say:

  • It does not say these robots are currently being used by all companies (though it mentions a lawsuit against one company, Workday).
  • It does not say we should stop using AI.
  • It does not say the "pro-Black" bias is a good thing; the authors argue that any bias (favoring one group over another) is bad for fairness. The goal is a neutral scale, not a tipped one.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →