Dress-ED: Instruction-Guided Editing for Virtual Try-On and Try-Off
This paper introduces Dress-ED, the first large-scale dataset unifying Virtual Try-On, Virtual Try-Off, and text-guided garment editing with over 146k instruction-driven samples, alongside a unified multimodal diffusion framework to enable controllable and interactive fashion generation.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you have a magical wardrobe in your living room. You pick out a dress, put it on, and instantly, you can change its color, turn it into leather, add a pocket, or even shorten the sleeves—all just by speaking to it. That's the dream of "Virtual Try-On" (VTON). But until now, this magic was a bit clumsy. You could try on a dress, but you couldn't easily edit it with specific instructions like "make the sleeves long" or "add a camouflage pattern."
This paper introduces Dress-ED, a massive new project that acts like a "training school" for AI to master this kind of fashion magic. Here is the breakdown in simple terms:
1. The Problem: The "Static Mannequin"
Think of previous fashion AI datasets as a library of mannequins. You could swap a mannequin's shirt for another one, but you couldn't tell the mannequin, "Hey, turn that blue shirt into a red leather jacket with a zipper." The AI didn't have a big enough library of examples showing how to make those specific changes while keeping the person looking like themselves.
2. The Solution: The "Fashion Chef" (Dress-ED)
The authors built Dress-ED, a giant cookbook with over 146,000 recipes.
The Ingredients: Every "recipe" (or data sample) has four parts:
- The original photo of a person wearing a garment.
- The original photo of the garment alone (like a catalog picture).
- The edited version of the garment (e.g., now it's yellow leather).
- The edited version of the person wearing that new garment.
- The Instruction: A text note saying exactly what changed (e.g., "Change the color to yellow").
The Magic Trick (How they made it): They didn't ask humans to draw 146,000 pictures (that would take forever). Instead, they built an automated assembly line:
- The Analyst (Qwen3-VL): A super-smart AI looks at a photo and writes a detailed description of the clothes (e.g., "Blue cotton shirt, short sleeves").
- The Chef (FLUX.2): Another AI takes that description and follows a command like "Make it yellow" to create a new, edited photo of the shirt.
- The Tailor (FitDiT): A third AI takes the new yellow shirt and "sews" it onto the person in the photo.
- The Inspector (GPT-5 & InternVL): A final AI acts as a strict quality control manager. It checks: "Did the shirt actually turn yellow? Did the person's face stay the same? Is the zipper in the right place?" If the answer is no, the sample is thrown in the trash. If yes, it goes into the dataset.
3. The Two Main Games
With this new dataset, they set up two challenges for AI models to solve:
- Virtual Try-On (The "Fitting Room"): You give the AI a person and a command ("Make the dress red"). The AI must show you the person wearing the red dress.
- Virtual Try-Off (The "Catalog"): You give the AI a person wearing a dress and a command ("Remove the belt"). The AI must generate a clean photo of just the dress without the belt, as if it were a product photo for a store.
4. The New Star: Dress-EM
The authors didn't just build the dataset; they also built a new AI model called Dress-EM to play these games.
- Think of Dress-EM as a bilingual fashion designer. It can read your English instructions ("Add a pocket") and look at the visual clues of the clothes at the same time.
- It uses a "connector" to bridge the gap between language and images, ensuring that when you ask for a change, it happens exactly where you want it, without messing up the rest of the outfit.
5. Why This Matters
Before this, if you wanted to edit clothes in a photo, you had to use Photoshop or hope a generic AI guessed what you meant.
- Dress-ED gives AI a massive, structured library of "before and after" examples with clear instructions.
- Dress-EM proves that AI can now understand complex fashion commands, like changing the material from cotton to leather or adding a specific pattern, while keeping the person's face and pose perfectly natural.
In a nutshell: The authors built a massive, automated "fashion school" where AI learns to listen to your voice and instantly redesign clothes on people. They created the curriculum (the dataset) and the top student (the model) to show that the future of online shopping might soon let you say, "I want this dress in green velvet with long sleeves," and see it happen instantly.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.