← Latest papers
💻 computer science

Face Anything: 4D Face Reconstruction from Any Image Sequence

This paper presents "Face Anything," a unified feed-forward transformer-based method that achieves state-of-the-art 4D facial reconstruction and tracking from any image sequence by predicting canonical facial coordinates to resolve geometric ambiguities and ensure temporal consistency.

Original authors: Umut Kocasari, Simon Giebenhain, Richard Shaw, Matthias Nießner

Published 2026-04-22
📖 4 min read☕ Coffee break read

Original authors: Umut Kocasari, Simon Giebenhain, Richard Shaw, Matthias Nießner

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to build a perfect, animated 3D movie of a person's face just by watching a video of them talking. This is incredibly hard because faces are tricky: they squish, stretch, and twist when we smile or frown, and the camera angle might change.

Most old methods tried to solve this by guessing how every single pixel moves from one frame to the next. It's like trying to track a specific drop of water in a rushing river; it's chaotic, confusing, and often leads to mistakes.

"Face Anything" is a new, smarter way to do this. Here is how it works, explained simply:

1. The "Universal Passport" Analogy

Instead of tracking how a pixel moves from Frame A to Frame B, this new method asks a different question: "Where does this pixel belong on the face's 'ID card'?"

Think of a human face like a globe. No matter how you turn your head or make a funny face, your nose is always in the center, and your eyes are always above it.

  • The Old Way: Trying to calculate the exact path a pixel takes as it flies across the screen.
  • The New Way (Face Anything): Assigning every pixel a "passport number" (a Canonical Coordinate). This passport tells the computer, "I am the pixel on the left cheek, 2 millimeters from the nose."

Because this "passport" is based on the face's structure rather than the camera's view, it stays the same even if the person turns their head or smiles wildly.

2. The "Master Blueprint"

The system builds a Master Blueprint (called the Canonical Space).

  • Imagine a clay model of a face that never moves.
  • When you take a photo of a real person, the AI instantly figures out: "Okay, this photo is just a distorted version of that Master Blueprint."
  • It maps every pixel in your photo back to its spot on the Master Blueprint.

Once everything is mapped to this Master Blueprint, connecting the dots between different moments in time becomes easy. You don't need to guess where the pixel went; you just look at its passport number. If the passport says "Left Eyebrow," you know exactly where that pixel is in the next frame, too.

3. The "Swiss Army Knife" Model

The paper introduces a single AI model (a "Swiss Army Knife") that does three jobs at once in one quick glance:

  1. Depth: It figures out how far away the face is (3D shape).
  2. Tracking: It follows every point on the face over time.
  3. Mapping: It assigns those "passport numbers" (canonical coordinates).

Because it does all this together, the results are much more consistent. The 3D shape doesn't jitter, and the tracking doesn't get lost when the person turns their head.

4. Why It's a Big Deal

  • Speed: It's like going from driving a car through traffic (old methods) to taking a teleporter (this method). It's roughly 3 times faster and much more accurate.
  • Stability: In old videos, 3D faces often look like they are "shaking" or melting. This method keeps the face solid and stable, like a high-quality statue that just happens to be talking.
  • The "Face Anything" Name: It's a play on the famous "Depth Anything" model. Just as that model could guess depth for anything in a picture, this model can reconstruct and track any face, no matter how wild the expression or lighting.

The Catch (Limitations)

The system is a specialist. It is an expert on faces.

  • If you point the camera at a face and a microphone, it will perfectly track the face but might get confused by the microphone because the microphone doesn't have a "face passport."
  • It relies on having learned what faces look like from millions of examples, so it struggles with things that aren't faces.

In a Nutshell

Face Anything is like giving every pixel on a face a permanent address. Instead of chasing pixels as they run around the screen, the computer just checks their address to know exactly where they are, what they look like, and how they connect to the rest of the face. This makes creating realistic 3D avatars from simple videos faster, smoother, and more accurate than ever before.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →