← Latest papers
💻 computer science

Measuring Validity in LLM-based Resume Screening

This paper addresses the challenge of evaluating LLM-based resume screening by introducing a novel dataset with known ground-truth candidate rankings, revealing that many models fail to consistently select more qualified candidates, struggle to abstain on equal qualifications, and exhibit demographic biases, thereby offering a principled framework for auditing these systems.

Original authors: Jane Castleman, Zeyu Shen, Blossom Metevier, Max Springer, Aleksandra Korolova

Published 2026-02-24
📖 5 min read🧠 Deep dive

Original authors: Jane Castleman, Zeyu Shen, Blossom Metevier, Max Springer, Aleksandra Korolova

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are hiring a new chef for a busy restaurant. You have thousands of applications, so you decide to use a super-smart robot (an AI) to read the resumes and pick the best candidates. You trust the robot because it's fast and knows a lot about cooking.

But here's the problem: How do you know the robot is actually picking the best chefs, or if it's just guessing, getting confused, or picking people based on their names instead of their cooking skills?

This paper is like a group of scientists building a giant, controlled cooking competition to test these robots. They didn't just ask the robots to judge real resumes (because real resumes are messy and we don't always know who the "perfect" chef is). Instead, they created a "lab" where they knew the answer beforehand.

Here is the breakdown of their experiment using simple analogies:

1. The "Fake Resume" Factory

The researchers realized that to test a robot, you need a test where you know the right answer. So, they built a factory that creates pairs of resumes:

  • The "Perfect" Pair: One resume has a chef who can bake a cake, and the other has a chef who can bake a cake and make a soufflé. The robot must pick the soufflé-maker. If it picks the cake-maker, it failed.
  • The "Twin" Pair: Two resumes are identical in every cooking skill. The only difference is that one says "John" and the other says "Jamal," or one mentions a "Women in Tech" award and the other doesn't. Since the cooking skills are the same, the robot should say, "I can't choose; they are equal." If it picks one over the other just because of the name or the award, it's being biased.

2. The Two Big Tests

The scientists gave these fake resume pairs to many different AI models (like the brains behind ChatGPT, Claude, and Gemini) to see how they performed. They measured two things:

A. The "Logic Test" (Criterion Validity)

  • The Question: Can the robot tell the difference between a good chef and a great chef?
  • The Result: Surprisingly, many robots struggled! Even the newest, most expensive models sometimes picked the less qualified candidate. It's like a robot that sometimes thinks a person who can only boil water is a better chef than someone who can bake a 5-course meal.
  • The "Abstain" Option: The robots were allowed to say, "I don't know, they are too close to call." When they were unsure, they often chose to say "I don't know" rather than picking the wrong person. This is actually a good thing! It's better for a robot to admit it's confused than to confidently hire the wrong person.

B. The "Fairness Test" (Discriminant Validity)

  • The Question: If two chefs are equally skilled, does the robot pick one just because of their name, gender, or race?
  • The Result: The robots were bad at this too. When two candidates were exactly equal, the robots often picked one anyway, as if they couldn't stand to leave the decision blank.
  • The Twist: In some cases, the robots seemed to be too careful about being fair. They sometimes picked candidates from historically marginalized groups (like Black women) at higher rates than White men, even when the candidates were identical. The researchers call this "Over-Alignment." It's like a referee who, trying so hard not to be racist, accidentally starts favoring one team so much that they ignore the actual rules of the game.

3. Why This Matters

The paper argues that we can't just assume these AI hiring tools work.

  • The "Black Box" Problem: Companies are using these tools "out of the box" without testing them. It's like buying a car without ever checking if the brakes work.
  • The Solution: The researchers created a blueprint for how to test these tools. Instead of trusting the AI company's word, anyone (like a government auditor or a small business) can use this blueprint to generate their own "fake resume" tests to see if the AI is actually doing its job.

The Big Takeaway

Think of these AI hiring tools as new drivers. Just because they have a license (they are smart and trained on lots of data) doesn't mean they are safe to drive on the highway.

  • Sometimes they miss obvious hazards (failing the Logic Test).
  • Sometimes they get distracted by things that don't matter, like the color of the other car (failing the Fairness Test).
  • Sometimes they try so hard to be "nice" that they drive the wrong way (Over-Alignment).

This paper gives us a driving test that we can use to make sure these AI drivers are actually ready for the road before they start hiring our future employees.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →