← Latest papers
💻 computer science

Position: Age Estimation Models Do Not Process Biometric Data

This position paper argues that age estimation models do not process biometric data because empirical evidence shows they lack the identity-discriminative capabilities required for individual identification, urging regulators to distinguish such transient processing from template storage to avoid unnecessary legal burdens.

Original authors: Nikita Marshalkin

Published 2026-05-19
📖 5 min read🧠 Deep dive

Original authors: Nikita Marshalkin

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

The Big Question: Is a "Guess" the Same as a "Fingerprint"?

Imagine you are walking into a store. A camera takes a picture of your face to guess your age. The question this paper asks is: Does that camera just "guess" your age, or does it secretly create a digital fingerprint that could identify you?

If the answer is "it creates a fingerprint," then strict privacy laws (like GDPR in Europe or BIPA in Illinois) kick in. These laws say you must give explicit permission, or the company could face huge fines. If the answer is "it just guesses," those strict rules might not apply.

The author, Nikita Marshalkin, argues that age estimation models do not process biometric data because they are terrible at identifying people. They are like a weather forecaster who is great at predicting rain but terrible at recognizing faces.

The Legal Confusion: "Can" vs. "Will"

The paper highlights a confusing gap in the law.

  • The "Purpose" Argument: Some regulators say, "If the company intends to guess your age and not identify you, it's fine."
  • The "Capability" Argument: Other regulators say, "If the technology could theoretically identify you, even if they don't want to, it's biometric data."

The paper focuses on the Capability Argument. It asks: Even if the company doesn't want to, does the math inside the computer actually allow them to identify you?

The Experiment: The "Taste Test"

To answer this, the author didn't just guess; they ran a massive experiment. Think of it like a taste test.

  1. The Setup: They took 14 different "age guesser" models (some commercial, some open-source).
  2. The Challenge: They asked these models a different question: "Look at these two photos. Are they the same person?"
    • They tested this on three different "obstacle courses" (benchmarks):
      • The Easy Course: Normal photos (LFW).
      • The Time Travel Course: Photos of the same person 30 years apart (AgeDB-30).
      • The Angle Course: One photo from the front, one from the side (CFP-FP).
  3. The Comparison: They compared the "age guessers" against a "face identifier" (a model built specifically to recognize people, like the one used for unlocking phones).

The Results: The Age Guessers Failed Hard

The results were clear and dramatic.

  • The Face Identifier: It was like a master detective. It got the answer right almost 100% of the time.
  • The Age Guessers: They were like a drunk person trying to recognize their own reflection. They failed miserably.

The Analogy:
Imagine a security checkpoint that requires a password to enter.

  • The Face Identifier has the correct password.
  • The Age Estimators are trying to guess the password by looking at the color of the sky. They are two orders of magnitude (100 times) worse than what is required to actually identify someone.

Even when the author tried to "teach" the age estimators how to recognize faces by adding a small training layer on top (like giving them a cheat sheet), they still failed on the harder tests. They simply didn't have the "ingredients" (data) inside them to build an identity.

The "Ghost in the Machine" Myth

Some people worry that even if the model doesn't store a face ID, it might create a temporary "ghost" (an intermediate representation) inside the computer that could be used to identify someone later.

The paper argues this is like worrying that a thermometer contains a secret map of your house.

  • A thermometer measures temperature. While it touches your skin, it creates a temporary reading.
  • Once it reads the temperature, it throws the reading away.
  • Even if you looked at the thermometer's internal gears, you couldn't reconstruct a map of your house from them.

Similarly, the paper shows that the "ghost" inside an age estimator is too blurry to identify anyone. It's not a fingerprint; it's a blurry smudge that only knows "this looks like a 25-year-old," not "this is John Doe."

The Call to Action

The paper ends with two requests:

  1. To Researchers: "Stop being vague." Tell the world clearly: Does this system save a face ID? No. Does it throw the data away immediately? Yes. Be transparent about what the system can and cannot do.
  2. To Regulators: "Stop treating all face cameras the same." Distinguish between a system that stores a face ID to find a person (like a police database) and a system that processes a face temporarily to guess an age and then deletes the data. The latter shouldn't be punished with the same heavy rules as the former.

The Bottom Line

The paper concludes that age estimation models cannot identify people. They are so bad at it that they fall far below the legal standards for what counts as "biometric data." Therefore, they shouldn't be forced to follow the strictest privacy rules designed for systems that actually know who you are.

Note: The author clarifies that this is not legal advice, but rather scientific evidence to help regulators make better decisions.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →