FashionPose: Unified Text-Driven Fashion Synthesis with Joint Geometric and Photometric Control
FashionPose is a unified, language-driven framework that overcomes the limitations of conventional pose-guided synthesis by employing a cascaded architecture to jointly generate template-free poses and scene-aware relighting, thereby enabling realistic and controllable garment synthesis for fashion e-commerce.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are a director on a movie set, but instead of hiring actors, you are commanding a digital artist to draw a person wearing a specific outfit. In the world of computer vision—the science of teaching computers to "see" and create images—this is a tricky job. Usually, if you want a computer to draw a person in a new pose, you have to give it a rigid skeleton, like a wireframe mannequin, to hold up the clothes. It's like trying to dress a puppet where you can only move the joints if you already have a pre-made blueprint. Furthermore, most of these digital artists work in a "studio" with flat, boring light. If you ask them to draw someone standing in a sunny park at sunset, the computer often gets confused, leaving the person looking like a flat sticker pasted onto a background that doesn't match the lighting. This paper tackles the problem of how to tell a computer to create a realistic fashion image using only your words, handling both the tricky body movements and the complex lighting without needing a pre-made skeleton or a studio setting.
The researchers behind this work, led by Chuancheng Shi and colleagues, propose a new system called FashionPose. Think of it as a "magic translator" that turns your natural language descriptions directly into a 3D-ready fashion show, skipping the need for rigid wireframes. They argue that the old way of doing things—relying on pre-defined skeletons and ignoring the environment—is too limiting. Instead, their system uses a three-step "cascaded" process to turn a simple sentence like "a girl standing by a sunny window in a cozy room" into a high-quality image where the girl's pose matches the words, and the sunlight hits her clothes exactly as described.
First, the system acts like a creative director who listens to your story and sketches a pose. Instead of looking up a pre-made skeleton, it uses a "bidirectional contrastive alignment" mechanism. Imagine this as a game of "hot and cold" where the computer learns to match the feeling of your words (like "crossed legs" or "arms raised") directly to the geometry of a body, without needing a template. To teach the computer how to do this, the team built a massive new library called PoseCap, containing over 40,000 pairs of text descriptions and body key points. This allows the AI to learn that "standing by a sunny window" isn't just a background detail; it's a clue for how the light should hit the person.
Next, the system moves to the "costume designer" phase. Once it has the pose, it generates a high-fidelity image of the person. Crucially, it uses an "identity-anchored" strategy. Imagine you have a photo of a specific model wearing a specific jacket. The system takes that jacket and the model's face, locks them in place, and then stretches and folds them into the new pose you described. This ensures that even if the person is doing a complex dance move, the buttons on the jacket and the model's face stay exactly the same, preventing the "melting face" or "distorted clothes" glitches common in other AI art tools.
Finally, the system handles the "lighting crew." This is where FashionPose shines compared to older methods. It includes a "prompt-conditioned relighting" module. If your text says "golden hour" or "soft studio lighting," this module adjusts the shadows and highlights on the generated person to match that specific atmosphere. It's like having a lighting technician who reads your script and moves the spotlights around to match the mood, ensuring the person doesn't look like they were cut out of a different photo and pasted in.
The team tested their system rigorously. They found that FashionPose outperformed existing methods in both accuracy and realism. In their tests, the system achieved a keypoint error (MSE) of 243.8, which is lower (better) than the next best methods which hovered around 262.2. When it came to the overall look of the images, measured by a metric called FID (where lower is better), FashionPose scored 5.6690, beating the previous best combination of tools which scored 6.3671. In user studies, where real people voted on which images looked best, FashionPose was preferred in 46.8% of cases for pose consistency and 55.9% for overall aesthetic quality.
The paper explicitly rules out the idea that you need a pre-existing skeleton or a "studio-like" neutral background to get good results. They argue that relying on external pose estimators limits what the AI can do and that ignoring the environment leads to unrealistic images where the lighting doesn't match the pose. By proving that a single text prompt can control both the geometry and the photometry (lighting) simultaneously, they suggest a new way forward for virtual fashion. While they don't claim to have solved every problem in the universe, their experiments suggest that this unified approach creates a much more robust and flexible solution for personalized, scene-aware virtual fashion displays, making it easier for anyone to visualize clothes in any setting just by typing a sentence.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.