← Latest papers
💻 computer science

Durian: Dual Reference Image-Guided Portrait Animation with Attribute Transfer

Durian is a novel portrait animation method that enables cross-identity attribute transfer from one or more reference images by employing a self-reconstruction training strategy with a Dual ReferenceNet and complementary masking to learn from unpaired video data, achieving state-of-the-art performance with support for multi-attribute composition and smooth interpolation.

Original authors: Hyunsoo Cha, Byungjun Kim, Hanbyul Joo

Published 2026-04-07
📖 5 min read🧠 Deep dive

Original authors: Hyunsoo Cha, Byungjun Kim, Hanbyul Joo

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have a photo of yourself, but you're bored with your current look. You want to try on a cool new hat, swap your hair for a wild afro, or add a pair of stylish sunglasses—all while you are talking, smiling, and moving your head naturally.

Usually, doing this is a nightmare for computers. To teach a computer how to do this, you'd typically need thousands of photos of the same person wearing every possible combination of hats, glasses, and hairstyles. That's impossible to collect.

Enter Durian (the new AI method from Seoul National University). Think of Durian as a magical, super-smart video editor that can learn to swap your look without ever needing a "before and after" photo of the same person.

Here is how it works, broken down into simple concepts:

1. The Problem: The "Missing Puzzle Piece"

Normally, AI learns by looking at pairs: "Here is Person A with short hair, and here is Person A with long hair." It learns to swap the hair. But we don't have millions of videos of the same person changing their hair every day.

The Durian Solution: Instead of waiting for perfect pairs, Durian learns by playing a game of "self-reconstruction" using random videos of people talking.

  • It grabs two random frames from the same video of a person.
  • It treats one frame as the "Target" (the person's face).
  • It treats the other frame as the "Reference" (the new hair or glasses).
  • It tries to rebuild the video, forcing itself to figure out: "How do I take the hair from Frame B and put it onto the face in Frame A, while keeping the face looking like Frame A?"

It's like learning to cook a new dish by tasting two different soups and guessing the recipe, rather than having the recipe book in front of you.

2. The Secret Sauce: The "Dual Reference" Kitchen

Most AI models have one brain to look at the face and one to look at the new hair. Durian has a Dual ReferenceNet, which is like a kitchen with two specialized chefs:

  • Chef Identity (PRNet): This chef only looks at the face and says, "I must keep the eyes, nose, and smile exactly the same."
  • Chef Attribute (ARNet): This chef only looks at the new hair or glasses and says, "I must copy this style perfectly."

These two chefs work separately but talk to each other constantly through a Spatial Attention mechanism. Think of this as a high-tech laser pointer. The laser tells the AI exactly where on the face the new hair should go and where the glasses should sit, ensuring they don't accidentally paste the hair over the eyes!

3. The Training Trick: "Masking" and "Stretching"

To make sure the AI doesn't get lazy (like just copying the whole face from the reference), Durian uses masks.

  • It covers up the hair on the reference photo so the AI only sees the hair.
  • It covers up the hair on the target photo so the AI must generate new hair there.

But here's the clever part: Mask Expansion.
Imagine you are teaching a child to draw a hat. If you only show them a small, tight hat, they might struggle if you ask them to draw a big floppy one later. Durian artificially "stretches" the training masks. It teaches the AI that a "hat" isn't just one specific shape; it can be big, small, tilted, or shifted. This makes the AI robust, meaning it won't break if the new glasses are slightly crooked or the new hair is wilder than the reference.

4. The Magic at the End: Mixing and Matching

Because Durian is so smart, it doesn't just do one thing at a time.

  • The "Smoothie" Effect (Interpolation): You can ask Durian to blend two hairstyles. It can create a video where your hair slowly morphs from a short bob to a long ponytail, frame by frame.
  • The "Combo Meal" (Multi-Attribute): You can give it a photo of a hat, a photo of glasses, and a photo of a beard. Durian can combine all three into a single video instantly. It knows that the hat goes over the hair and the glasses go on the nose, handling the overlaps perfectly.

Why is this a big deal?

Before Durian, if you wanted to try on a new hairstyle in a video, you needed a 3D scanner, a specific app, or a pre-made template that looked stiff and fake.

Durian is the "Zero-Shot" Wizard.

  • Zero-Shot: It doesn't need to be retrained for every new person or new object.
  • In-the-Wild: It works with random photos and videos you find on the internet.
  • One-Pass: It does the swapping and the animation in one go, like a single magic spell, rather than a long, complicated process.

In short: Durian is the first AI that can look at a photo of a stranger's cool haircut and a video of you, and seamlessly edit your video to have that haircut, all while you talk and move, without ever needing a dataset of you changing your hair. It's like having a personal stylist who can instantly try on any look from any photo, right on your face.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →