Why Aggregate Accuracy is Inadequate for Evaluating Fairness in Law Enforcement Facial Recognition Systems
This paper argues that aggregate accuracy is an insufficient metric for evaluating law enforcement facial recognition systems because it obscures critical demographic disparities in error rates, necessitating a shift toward fairness-aware, subgroup-level evaluation frameworks to prevent disproportionate societal harm.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are the captain of a ship, and you have a new, high-tech radar system to help you spot other ships in the fog. The manufacturer tells you, "This radar is 90% accurate!" That sounds amazing, right? You feel safe.
But here's the catch: The manufacturer didn't tell you who the radar misses.
This paper, written by Khalid Adnan Alsayed, argues that in the world of law enforcement, using a single "accuracy score" to judge facial recognition software is like judging that radar system only by its overall success rate. It's a dangerous game because it hides a terrifying secret: The system might be working perfectly for some people, but failing miserably for others.
Here is a breakdown of the paper's main points using simple analogies:
1. The "Average" Trap
The Analogy: Imagine a classroom where the teacher asks, "How did the class do on the test?" The teacher says, "The average score was 85%!" That sounds great. But what if the top 50% of students got 100%, and the bottom 50% got 0%? The "average" hides the fact that half the class failed completely.
The Reality: Facial recognition systems often have a high "aggregate accuracy" (like that 85% average). But when you look closer, the system might be 99% accurate for light-skinned men but only 60% accurate for dark-skinned women. The high overall score makes the system look fair and reliable, but it's actually hiding a massive gap in performance.
2. The Two Types of Mistakes (False Positives vs. False Negatives)
In law enforcement, getting a number wrong isn't just a math error; it changes lives. The paper highlights two specific ways the system can fail, and why the "average" doesn't tell us which group is suffering:
- False Positive (The "Wrong Guy" Alarm): The system looks at an innocent person and says, "That's the criminal!"
- The Analogy: It's like a smoke detector that goes off when you just toasted a piece of bread. If this happens too often for a specific group of people, they get stopped, questioned, or even arrested for crimes they didn't commit.
- False Negative (The "Missed Guy"): The system looks at a real criminal and says, "I don't know who that is."
- The Analogy: It's like a security guard who misses a thief walking right past them. If this happens more often for a specific group, those criminals might go free, or investigations might stall.
The Problem: If you only look at the "total accuracy," you can't see that Group A is getting falsely accused (False Positives) while Group B is getting missed entirely (False Negatives).
3. The "Black Box" Problem
The Analogy: Imagine you buy a car from a mysterious company. They say, "It's the fastest car in the world!" But they won't let you open the hood to see the engine, and they won't tell you how the brakes work. You just have to trust them.
The Reality: Many police departments buy facial recognition software from private companies. These companies treat their code as a "trade secret" (a black box). Because the police can't see inside the code to check for bias, they rely on the simple "accuracy number" the company gives them. This paper argues we need a way to "audit" these cars from the outside—checking how they drive in real life without needing to see the engine.
4. The Tough Choice: Fairness vs. Perfection
The Analogy: Imagine you are baking a cake for a huge party. You want the cake to be perfect for everyone. But you realize that the recipe works great for people who like chocolate, but it's too dry for people who like vanilla.
- Option A: Keep the recipe exactly as is. The chocolate lovers get a perfect cake, but the vanilla lovers get a dry, bad cake. The average taste score is high.
- Option B: Change the recipe to make it "good enough" for everyone. The chocolate cake isn't quite as perfect as before, but now the vanilla lovers can actually enjoy it too. The average taste score might drop slightly, but everyone is happier.
The Reality: The paper admits that making a system fairer might slightly lower its overall accuracy. But in law enforcement, a "slightly less accurate" system that doesn't unfairly target specific groups is better than a "perfectly accurate" system that ruins innocent lives.
The Bottom Line
This paper is a wake-up call. It says: "Stop looking at the headline number."
Just because a facial recognition system says it is "95% accurate" doesn't mean it is safe to use in a police station. We need to look under the hood and ask:
- Who is it getting wrong?
- Is it accusing innocent people from certain groups more often?
- Is it missing criminals from certain groups more often?
Until we start measuring fairness (how errors are shared among different groups) instead of just accuracy (the total score), we risk building a justice system that is mathematically "smart" but socially unjust.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.