← Latest papers
💻 computer science

FG-Portrait: 3D Flow Guided Editable Portrait Animation

FG-Portrait introduces a 3D flow-guided diffusion framework that leverages geometry-driven 3D flows and depth-guided sampling to achieve high-fidelity, editable portrait animation with superior motion transfer and identity preservation compared to existing methods.

Original authors: Yating Xu, Yunqi Miao, Evangelos Ververas, Jiankang Deng, Jifei Song

Published 2026-03-25
📖 4 min read☕ Coffee break read

Original authors: Yating Xu, Yunqi Miao, Evangelos Ververas, Jiankang Deng, Jifei Song

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have a photo of your friend, Alice, sitting still and smiling. Now, imagine you have a video of your other friend, Bob, making funny faces, tilting his head, and talking.

The Goal: You want to create a new video where Alice is the one making all those funny faces and moving her head exactly like Bob, but she still looks exactly like Alice.

This is called Portrait Animation. It's like a digital puppet show, but instead of strings, we use math and AI.

The Problem with Old Methods

Think of old AI methods as a clumsy puppeteer. They try to guess how Alice's face should move just by looking at Bob's video.

  • The Mistake: They often get confused. If Bob tilts his head left, the AI might accidentally make Alice's nose look like it's melting or her eyes might get swapped. It's like trying to copy a dance move by only watching the feet, without understanding how the whole body connects.
  • The Result: The movement looks "off," or Alice starts to look a little bit like Bob.

The New Solution: FG-Portrait (The "3D Blueprint" Approach)

The authors of this paper, FG-Portrait, came up with a smarter way. Instead of guessing, they use a 3D Blueprint.

Here is the analogy:

  1. The Invisible Mannequin: Imagine that under every photo of a human face, there is a perfect, invisible 3D mannequin (a digital clay model). This mannequin knows exactly where every nose, cheek, and chin is in 3D space.
  2. The "Flow" Map: When Bob moves his head, his invisible mannequin moves. When Alice sits still, her mannequin is still.
    • The new method draws a tiny, invisible string (a "flow") from every single point on Bob's moving mannequin back to the matching point on Alice's still mannequin.
    • Analogy: It's like having a GPS map that says, "If Bob's left cheek moves here, Alice's left cheek must have come from there."
  3. The Magic Translation: The AI uses these invisible strings to tell the computer: "Don't guess! Just move Alice's pixels exactly where the 3D map says they should go."

Why is this better?

  • No Guessing: Because the 3D mannequin is a perfect mathematical model, the AI doesn't have to guess how a nose moves when a head turns. It knows the geometry perfectly.
  • The "Depth" Trick: Sometimes, a 2D photo is tricky (is that a shadow or a deep wrinkle?). The paper adds a "Depth-Guided" step. Think of this as the AI putting on 3D glasses to understand exactly how far away a point is, ensuring the movement strings are pulled from the right spot.

The Cool Extra Feature: "Remote Control" Editing

The best part? This system isn't just for copying Bob.

  • Once the AI understands the 3D structure, you can act as the director.
  • You can tell the AI: "Make Alice smile more than Bob did," or "Make her tilt her head less."
  • It's like having a remote control for the 3D mannequin. You can tweak the expression or the pose after the animation is generated, and the AI will adjust the picture instantly.

Summary in a Nutshell

  • Old Way: "I think if Bob moves his head left, Alice should move her pixels left." (Often results in a messy, melted face).
  • FG-Portrait Way: "I have a 3D map of Bob and Alice. I know exactly which of Alice's pixels corresponds to Bob's moving nose. I will move them precisely." (Results in a crisp, realistic, and accurate animation).

This method makes digital avatars look much more real and allows us to edit their expressions and poses with the precision of a 3D sculptor, rather than a clumsy painter.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →