Beyond Facial Consistency: Personalized Person Image Generation with Holistic Identity Preservation
This paper introduces a dual-branch framework enhanced by a Dynamic Balancing Scaling (DBS) strategy to resolve the trade-off between facial fidelity and overall appearance consistency in personalized person image generation, alongside the release of the Pexels-100 benchmark for holistic identity evaluation.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are a director in a movie studio, but instead of hiring actors, you have a magical camera that can conjure up any character you can describe. You want a hero who looks exactly like your best friend—same smile, same eyes, same quirky nose—but wearing a space suit on Mars, or perhaps riding a dragon in a medieval forest. This is the dream of "personalized image generation." For a long time, AI artists had to choose between two frustrating options: either make the character look exactly like your friend's face, but forget what they were wearing or how their body looked; or dress them in the perfect outfit and pose, but end up with a stranger's face. It was like trying to bake a cake where you could get the frosting perfect or the sponge perfect, but never both at the same time. This paper dives into that specific corner of artificial intelligence, tackling the tricky problem of keeping a person's "whole self" consistent—their face and their entire outfit and vibe—when the AI is asked to put them in new situations.
The researchers behind this study, led by Yuxuan Xiao and colleagues, decided to stop choosing sides. They started by building a "Naive Dual-Branch" system, which is essentially a two-lane highway for the AI. One lane focuses on the big picture (the clothes, the body shape, the hairstyle), and the other lane zooms in on the tiny details of the face. They found that while this two-lane approach was a good start, it was a bit chaotic. The "face lane" was driving way too fast and taking over the whole car, while the "appearance lane" was getting lost in the rearview mirror. The result? Great faces, but the clothes would sometimes morph into something weird, or the body shape would get squished.
To fix this traffic jam, the team invented a clever traffic control system called Dynamic Balancing Scaling (DBS). Think of it as a smart conductor for an orchestra. The conductor uses two main tricks. First, they use Adaptive Temporal Gating, which is like a dimmer switch that changes the volume of the face and body musicians at different moments during the song. Early in the process, the conductor might let the face play a steady, reliable tune to set the identity, but as the song progresses, they gently turn up the volume on the body and clothes so they don't get drowned out. Second, they introduced Region-Aware Optimization. Imagine the AI is painting a picture, and it usually paints the whole canvas with the same brush pressure. This new trick tells the AI, "Hey, be super careful and precise when painting the face, but don't forget to give the clothes and background just as much attention." This ensures the AI doesn't obsess over the eyes and ignore the shirt.
The team tested their new method on a custom benchmark they created called Pexels-100, which features 100 different full-body images to see how well the AI could keep a person's identity intact across various scenes. The results were promising. Their method managed to strike a much better balance than other open-source tools, keeping the face recognizable while also making sure the clothes and body looked right. In fact, their approach performed so well that it even rivaled some expensive, closed-source commercial systems. However, the authors are careful to note that this isn't a magic wand for every problem yet. If you ask the AI to draw a person doing a complex gymnastic flip or showing intricate hand movements, the image might still get a little wobbly or distorted, because the underlying AI model still struggles with those tricky physical structures. But for most everyday scenarios, this new "conductor" helps the AI finally sing in harmony, keeping the whole person looking like themselves, not just their face.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.