← Latest papers
💻 computer science

Agent-Driven Multi-View Facial Prior Learning for Identity-Preserving Pose-Guided Generation

The paper proposes ADF-Pose, an agent-driven framework featuring an Identity Prior Extractor and a Multimodal Identity Converter to effectively preserve facial identity and fine-grained details in pose-guided person image generation, particularly under large pose variations where existing methods struggle.

Original authors: Zhaoyang Liu, Luosi Yan, Mubai Li, Xu Zheng, Haotian Yang

Published 2026-08-28
📖 5 min read🧠 Deep dive

Original authors: Zhaoyang Liu, Luosi Yan, Mubai Li, Xu Zheng, Haotian Yang

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine trying to take a photograph of a friend, but instead of asking them to turn their head, you are asking a computer to imagine what they would look like if they were facing a different direction. The goal is to create a new picture that looks exactly like that person, keeping their unique face, the shape of their jaw, and the texture of their skin, while placing them in a new pose. This is the challenge of pose-guided image generation. For years, computers have become quite good at moving a person's body, shifting their arms and legs to match a new stance. However, when the computer tries to turn the head, the result often feels wrong. The face might lose its specific shape, the eyes might drift apart, or the skin might look like a smooth, plastic mask. The computer struggles because it usually looks at the whole person as one big picture, missing the tiny, crucial details that make a face truly unique, especially when that face is turned away from the camera.

A team of researchers has developed a new approach to solve this problem, creating a system that treats the face with a level of care it has never received before. Instead of relying on a single, blurry guess of what the face might look like from a new angle, their system builds a detailed mental model of the person's face from multiple viewpoints. They call their method ADF-Pose. It works by first isolating the face in the original photo and then using a smart, automated assistant to imagine the missing sides of the face. If the person in the photo is looking slightly to the left, the system doesn't just guess what the right side looks like; it constructs a logical, verified version of that side, checking its work to ensure the nose, mouth, and jawline still belong to the same person. This process creates a set of "facial priors," which are essentially verified blueprints of the person's face from the front, left, and right.

Once these blueprints are ready, the system combines them with a written description of the person's facial features. It then feeds this combined information into a powerful image generator, which is a type of artificial intelligence known for creating realistic pictures from scratch. The generator uses these detailed blueprints and descriptions at every single step of the drawing process, from the initial rough sketch to the final fine-tuning of skin texture. This ensures that the identity of the person is locked in place, preventing the face from drifting or changing shape as the body moves. The researchers tested this method on large collections of fashion and street photography images, comparing their results against the best existing techniques. They found that their system was significantly better at keeping the face looking like the original person, even when the head was turned sharply to the side.

The results showed a clear improvement in how well the system preserved the person's identity. In tests using a dataset of fashion images, the new method achieved a score of 0.312 for facial similarity, a noticeable jump from the previous best score of 0.275. It also produced images with sharper details, reducing the blurriness of facial textures by about 7.4 percent compared to the leading alternative. When looking at the images side by side, the difference is striking. Older methods often produced faces that looked plausible but generic, with softened features or slightly wrong proportions. The new system, however, kept the specific contours of the jaw, the exact spacing of the eyes, and the fine lines around the mouth, even when the person was rotated to a profile view. The system managed to maintain these details without sacrificing the realism of the rest of the body, keeping the clothing and posture looking natural.

The researchers also broke down their system to understand which parts were doing the heavy lifting. They found that the step where the system constructs and verifies the missing views of the face was the most critical. Without this step, the face would lose its unique shape. The second most important part was the way the system combined the visual blueprints with the written descriptions, ensuring the computer understood not just what the face looked like, but how its features related to each other. Finally, the method of feeding this information into the image generator at every stage of the process ensured that the identity details were not lost as the image became clearer. While the system is not perfect and can still struggle if the original face is heavily blocked or too small to see clearly, it represents a significant step forward. It moves beyond simply moving pixels around to actually understanding and preserving the complex, unique geometry of a human face, allowing computers to generate new poses that feel truly like the person they are meant to represent.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →