← Latest papers
🤖 AI

FaithfulFaces: Pose-Faithful Facial Identity Preservation for Text-to-Video Generation

FaithfulFaces is a novel framework for text-to-video generation that addresses identity distortion in complex dynamic scenes by introducing a pose-shared identity aligner and a curated diverse dataset to maintain robust facial identity consistency across large pose variations and occlusions.

Original authors: Yuanzhi Wang, Xuhua Ren, Jiaxiang Cheng, Bing Ma, Kai Yu, Sen Liang, Wenyue Li, Tianxiang Zheng, Qinglin Lu, Zhen Cui

Published 2026-05-07
📖 4 min read☕ Coffee break read

Original authors: Yuanzhi Wang, Xuhua Ren, Jiaxiang Cheng, Bing Ma, Kai Yu, Sen Liang, Wenyue Li, Tianxiang Zheng, Qinglin Lu, Zhen Cui

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you want to create a short movie where a specific person (let's call him "Bob") is boxing. You have a single photo of Bob, and you want the AI to generate a video of him punching, dodging, and moving around.

The problem with current AI tools is that they are like amnesia patients. When Bob turns his head to the left, the AI forgets what Bob looks like from that angle. It might suddenly turn him into a different person, or his face might melt into a weird blob. If he puts his hand in front of his face, the AI panics and distorts his features.

The paper "FaithfulFaces" introduces a new system designed to solve this. Here is how it works, explained simply:

1. The Core Problem: The "Single-View" Trap

Most existing AI models only look at your reference photo and say, "Okay, I see Bob's face straight on." They don't understand that Bob's face is a 3D object that can turn, tilt, and twist. When the video asks for a side profile, the AI tries to guess, and it usually guesses wrong, leading to a distorted face.

2. The Solution: A "Pose-Faithful" Translator

The authors built a new framework called FaithfulFaces. Think of this framework as a super-smart translator that sits between your photo and the video generator.

  • The Old Way: The translator just says, "Here is Bob's face."
  • The FaithfulFaces Way: The translator says, "Here is Bob's face, and here is exactly how his head is tilted, turned, and rotated in 3D space."

3. How It Works: The "Pose Dictionary"

The secret sauce of FaithfulFaces is a component called the Pose-Shared Identity Aligner.

Imagine a giant library (a dictionary) filled with cards.

  • The Old Library: Had cards for "Bob's face" but no cards for "Bob's face turned left" or "Bob's face looking up."
  • The FaithfulFaces Library: Has a special Pose Dictionary. It learns that "Bob's face" is the same person whether he is looking left, right, up, or down.

When the system sees your photo, it doesn't just copy the pixels. It:

  1. Measures the Pose: It calculates the exact angle of the head (like using a 3D protractor to measure the tilt).
  2. Consults the Dictionary: It looks up how "Bob" should look at that specific angle.
  3. Aligns the Identity: It ensures that even if the angle changes, the "soul" of the face (the identity) stays consistent.

4. The "Training Gym"

To teach this system, the researchers couldn't just use random videos. They needed videos where people were moving their heads wildly.

  • They built a special dataset pipeline (a factory line) that hunted down thousands of videos.
  • They filtered out boring videos where people stood still.
  • They kept only the videos where people were boxing, dancing, or turning their heads significantly.
  • They fed these "wild movement" videos to the AI so it could learn: "Okay, when the head turns 90 degrees, the nose moves here, but the eyes still look like Bob's."

5. The Result: A Faithful Actor

In their tests, they asked the AI to generate a video of a person boxing (a scenario full of fast head movements and hands blocking the face).

  • Competitors (like Kling or ConsisID): The boxer's face would morph, look like a different person, or lose its shape when he turned his head.
  • FaithfulFaces: The boxer kept looking like the same person, with clear features, even when he was punching, dodging, or looking up.

Summary Analogy

If making a video with current AI is like trying to draw a person from a single photo and then guessing what they look like from the side (often resulting in a caricature), FaithfulFaces is like having a 3D sculptor who knows exactly how the person's face is built. No matter how the person moves, the sculptor ensures the clay retains the correct shape and identity.

The paper claims this method is the best at keeping the person's face looking real and consistent, even when they are doing complex, dynamic actions.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →