Oxygen-TryOn: Fashion-Native Foundation Model for Any-item Virtual Try-On
Oxygen-TryOn is a unified, fashion-native foundation model that achieves state-of-the-art photorealistic virtual try-on for diverse items and scenarios by leveraging a dedicated data engine and a three-stage training pipeline involving reinforcement learning, outperforming both leading proprietary and open-source systems.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine a world where you could instantly try on any outfit you see in a magazine, a friend's photo, or a random street snap, without ever stepping foot in a dressing room. This is the dream of "virtual try-on," a branch of artificial intelligence that tries to digitally dress people in new clothes. For a long time, these digital tailors were like clumsy apprentices: they could only handle one type of shirt at a time, they often forgot what the person looked like underneath, and they struggled to make the fabric look real when the person moved. They treated the task like a simple "paint-by-numbers" game, just filling in a blank space with a new texture. But what if the AI could actually understand fashion? What if it knew that a hat goes on a head, a bag hangs on a shoulder, and that a jacket needs to drape over a sweater, not just sit on top of it? This is the question researchers are tackling: how to build an AI that doesn't just paste images together, but truly "gets" how clothes, accessories, and people interact in the real world.
Enter Oxygen-TryOn, a new "fashion-native" foundation model created by the Oxygen AIGC Group and Joy Future Academy at JD. Think of this model not as a generic photo editor that you have to trick into doing fashion work, but as a specialist tailor who was born and raised in the world of style. While other AI systems try to force a general-purpose image editor to do virtual try-on by giving it vague instructions (which often leads to hallucinations, like inventing fake buttons or losing the person's face), Oxygen-TryOn was built from the ground up specifically for dressing people. It treats try-on not as a simple "fill-in-the-blank" puzzle, but as a complex reasoning task where the AI must figure out what each item is, where it belongs, and how it should look when worn.
The researchers found that by feeding this model a massive, specially curated library of fashion data and teaching it through a three-step "training recipe," they could achieve results that are far more realistic and consistent than previous methods. The model can take a single photo of a person and dress them in a single item, or it can juggle a chaotic mix of references—like a shirt from one picture, shoes from another, a bag from a third, and a hat from a fourth—and combine them into a single, coherent outfit. It handles everything from clean product photos to "in-the-wild" snapshots of people already wearing the items. Crucially, it preserves the person's identity (their face, body shape, and pose) and the fine details of the clothes (textures, logos, and patterns) with high fidelity.
In their experiments, Oxygen-TryOn didn't just play nice; it outperformed the competition. On standard tests, it achieved state-of-the-art scores, beating both powerful open-source models and top-tier proprietary systems like Nano Banana Pro, GPT-Image-2, and Seedream5 Lite. It particularly shines in "multi-item" scenarios, where it successfully layers multiple items without getting confused or dropping details. The team also showed that the model is flexible enough to follow extra instructions, like changing a person's pose or the background, all in the same generation pass.
The secret sauce behind this success wasn't just the model architecture, but the "data engine" the team built to feed it. They didn't just scrape the internet; they constructed a pipeline that collects, filters, annotates, and manufactures high-quality try-on data at a massive scale. They then trained the model in three distinct stages: first, a "continued pre-training" to learn the basics of fashion; second, a "supervised fine-tuning" to master the specific task of dressing; and finally, a "reinforcement learning" stage. This last step is like having a strict fashion critic (a hybrid reward system) grade the model's work, teaching it to fix subtle errors like awkward folds or identity drift. The result is a system that can handle diverse scenarios, from full-body outfits to half-body accessories, and even works on non-human subjects like anime characters or statues, proving that it has learned the general rules of wearing things rather than just memorizing specific photos.
In short, Oxygen-TryOn represents a shift from "guessing" what clothes should look like to "reasoning" about how they fit. It suggests that by building models specifically for a domain and training them with high-quality, structured data, we can create AI that is not only more accurate but also more useful for real-world applications like online shopping. While the model currently has limits—such as handling more than four reference items at once due to its underlying architecture—the paper demonstrates that the path to truly universal virtual try-on is open, and the future of digital fashion might just be a single, smart generation away.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.