← Latest papers
🤖 AI

AI-Powered Facial Mask Removal Is Not Suitable For Biometric Identification

This paper presents a large-scale analysis demonstrating that AI-powered facial mask removal, often used in crowd-sourced investigations, fails to produce reliable biometric matches and poses significant risks of misidentification.

Original authors: Emily A Cooper, Hany Farid

Published 2026-03-31
📖 5 min read🧠 Deep dive

Original authors: Emily A Cooper, Hany Farid

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

🕵️‍♀️ The Big Idea: Why "AI Unmasking" is a Bad Idea for Police Work

Imagine you are trying to identify a suspect in a crime. You have a grainy photo of their face, but they are wearing a mask or the photo is blurry. Someone suggests, "Let's use AI to guess what their face looks like underneath!"

This paper says: Stop. That is a terrible idea.

The authors, researchers from UC Berkeley, ran a massive experiment to prove that when you use AI to "fill in the blanks" of a masked face, the result is not the real person. It's more like a creative artist's guess than a forensic fact. Using these AI-generated faces to identify someone could lead to innocent people being wrongly accused.


🎨 The Experiment: The "Fill-in-the-Blanks" Game

To test this, the researchers played a game with four different groups of faces:

  1. Strangers: 400 random people.
  2. Twins: 100 people, each with two photos (same person, different day).
  3. Politicians: 91 US Senators wearing masks in one photo and no mask in another.
  4. Lookalikes: 63 pairs of celebrities who look very similar but aren't related.

They took these faces, covered the bottom half (like a mask), and asked three famous AI tools (ChatGPT, Gemini, and GrokAI) to "draw" the missing mouth and chin.

The Analogy:
Think of it like a "Connect the Dots" puzzle where half the dots are missing.

  • The Real Goal: You want the AI to draw the exact missing dots so the picture matches the original person perfectly.
  • What Actually Happened: The AI didn't look up the person's file. Instead, it looked at millions of other faces in its training data and said, "Hmm, most people with this nose shape have a mouth like this." It drew a statistically average mouth, not the specific mouth of the person in the photo.

📊 The Results: The "Guess" vs. The "Real Thing"

The researchers used a high-tech "face scanner" (a biometric system) to compare the Original Photo against the AI-Generated Photo.

Here is how they scored it (on a scale of -1 to 1, where 1 is a perfect match):

  • Same Person (Real Photos): Score of 0.71. (This is what we expect when comparing a person to themselves).
  • Total Strangers: Score of 0.04. (These look nothing alike).
  • Famous Lookalikes: Score of 0.15. (They look a bit similar, but not the same).
  • AI-Generated Faces: The scores landed right in the middle, around 0.34 to 0.52.

The Takeaway:
The AI-generated faces were not the real person. They were better than a random stranger, but they were nowhere near close enough to be the same person.

The Analogy:
Imagine you are trying to identify your best friend, "Bob."

  • Real Bob: You see him, and you say, "That's Bob!" (100% confidence).
  • AI Bob: The AI draws a picture of a guy who looks a bit like Bob. He has Bob's nose and hair, but the mouth is wrong. You might say, "Hey, that looks like Bob's cousin," but you wouldn't arrest that guy thinking it's Bob.
  • The Danger: If a crowd of people on the internet sees the AI drawing and says, "That's definitely Bob!" they might start accusing the wrong person.

🎲 The "Roll of the Dice" Problem

The researchers also discovered that AI is non-deterministic. This is a fancy way of saying: If you ask the AI the same question twice, it gives you two different answers.

They took one photo of a Senator and asked ChatGPT to unmask it 100 times.

  • Result: They got 100 different faces.
  • Some looked a little like the Senator; some looked very different.
  • The Analogy: It's like asking a chef to "make a burger" 100 times. Sometimes they give you a cheeseburger, sometimes a veggie burger, sometimes a burger with extra pickles. They are all "burgers," but they aren't the same burger. You can't use a random burger to prove who ordered it.

⚖️ Why This Matters (The "GrokAI" Incident)

The paper mentions a real-world disaster that sparked this research. A masked ICE agent killed a civilian. Someone used AI (GrokAI) to remove the mask. The AI drew a face that looked like a random guy named "Steve Grove."

Because the AI's guess looked "realistic" enough, people on social media convinced themselves that Steve Grove was the killer. It was a complete misidentification caused by the AI hallucinating a face that didn't exist.

🏁 The Final Verdict

Generative AI is a great artist, but a terrible detective.

  • What it does well: It creates beautiful, realistic-looking images based on what is likely to be there.
  • What it fails at: It cannot reconstruct the exact truth of a specific person's face.

The Bottom Line:
If you see an "AI-unmasked" photo of a criminal, do not trust it. It is a work of fiction, not a piece of evidence. Using these images for biometric identification (like matching a face to a database) is scientifically unsound and dangerous.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →