← Latest papers
🤖 AI

Demographic Fairness in Multimodal LLMs: A Benchmark of Gender and Ethnicity Bias in Face Verification

This paper presents a benchmarking study evaluating the gender and ethnicity bias of nine open-source Multimodal Large Language Models in face verification tasks, revealing that specialized models outperform general-purpose ones while demonstrating that higher accuracy does not guarantee demographic fairness and that bias patterns differ significantly from traditional face recognition systems.

Original authors: Ünsal Öztürk, Hatef Otroshi Shahreza, Sébastien Marcel

Published 2026-03-27
📖 5 min read🧠 Deep dive

Original authors: Ünsal Öztürk, Hatef Otroshi Shahreza, Sébastien Marcel

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have a new, super-smart robot assistant that can look at pictures and answer questions about them. You might ask it, "Is this person the same as that person?" This is called face verification.

For years, we've used specialized "face detectives" (traditional systems) for this job. They are like master fingerprint analysts: they turn a face into a secret code and compare the codes. But recently, we've started using these giant, general-purpose "brainy robots" (called Multimodal Large Language Models or MLLMs) to do the same thing. Instead of using secret codes, you just show them two photos and ask, "Are these the same person?" and they give you a score.

This paper is a big report card on how fair these new "brainy robots" are when they try to identify people from different backgrounds.

The Big Question: Are They Fair?

We know that old-school face detectives sometimes get it wrong more often for people with darker skin or specific ethnicities. It's like a security guard who is great at spotting one type of face but keeps misidentifying others.

The researchers wanted to know: Do these new, general-purpose AI robots have the same bias, or is it different?

The Experiment: A Test Drive

The researchers took 9 different AI models (ranging from small, 2-billion-parameter "puppies" to larger, 8-billion-parameter "dogs") and put them through a driving test.

  • The Course: They used two famous driving tracks (datasets): IJB-C (a huge, messy track with lots of different people) and RFW (a track specifically designed to test racial bias).
  • The Drivers: They tested how well these models could tell if two photos were the same person, breaking the results down by Ethnicity (African, East Asian, South Asian, Caucasian) and Gender (Male, Female).
  • The Scorecard: They didn't just look at who got the most answers right (accuracy). They also looked at Fairness. Did the robot make mistakes equally for everyone, or did it fail more often for specific groups?

The Surprising Results

1. The "Specialist" vs. The "Generalist"
There was one model in the test called FaceLLM-8B. Think of this model as a specialist detective who only studies faces. The other eight models were generalists (like a Swiss Army knife) that can do many things but aren't experts at faces.

  • Result: The specialist detective (FaceLLM) was way better at the job than the generalists. It made far fewer mistakes.
  • The Catch: Even the specialist wasn't as good as the old-school "code-based" systems. The generalist robots were often just guessing, performing no better than flipping a coin.

2. The Bias Flip-Flop
Here is the most interesting part. In traditional systems, people with darker skin often get the worst results.

  • The Twist: On the big, messy track (IJB-C), the new AI models actually struggled more with Caucasian faces than with others in many cases! It's like a security guard who is usually great at spotting everyone but suddenly gets confused by the people they know best.
  • On the other track (RFW): The bias flipped back. The models struggled most with African faces, which is what we usually expect.
  • Lesson: You can't assume these new AIs will have the same biases as the old ones. They are unpredictable.

3. The "Bad but Fair" Trap
The researchers found a tricky situation. Sometimes, a model is so bad at its job that it fails everyone equally.

  • Analogy: Imagine a terrible archer who misses the target every single time, no matter who is standing there. If they miss everyone equally, they look "fair" because the error rate is the same for all groups.
  • Reality: Being "fair" when you are useless isn't actually good! The most accurate models (like the specialist) weren't always the fairest, and the "fair" ones were often just the ones that were too dumb to get it right for anyone.

4. Gender vs. Race
The models were generally much fairer regarding gender (Male vs. Female) than they were regarding ethnicity. The gap between how well they treated men and women was tiny, but the gap between how they treated different ethnic groups was much wider.

The Bottom Line

This paper tells us that while these new "brainy robots" are exciting, we can't just trust them to do face verification yet.

  • They aren't as accurate as the old, specialized systems.
  • Their biases are weird and different from what we are used to (sometimes they mess up the "majority" group more!).
  • A model can look "fair" just because it's bad at its job.

The Takeaway: If we want to use these powerful new tools for security or ID checks, we need to be very careful. We can't just assume they are neutral. We need to test them constantly, just like we test a new car before letting it drive on the highway.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →