Grounding Free-Form Instructions for Fashion Complementary Image Generation
This paper introduces a new multimodal grounding task for fashion complementary image generation using free-form natural language instructions, supported by enriched benchmarks and the proposed StyleFlow model, which effectively generates stylistically coherent garments while reducing architectural complexity compared to existing approaches.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
In the world of fashion, the challenge has always been about more than just finding a single item that looks good; it is about assembling a complete look where every piece belongs together. For decades, computer scientists have tried to teach machines to understand this sense of style, a task known as complementary image generation. The goal is simple in theory: show the computer a shirt, and have it invent a pair of pants that matches it perfectly. Early attempts relied on rigid rules or simple templates, asking the machine to produce "a photo of jeans" without much room for nuance. However, real people do not speak in rigid templates. When a shopper searches for an outfit, they might ask for "relaxed-fit gym sweatpants with an elastic waistband" or simply "jeans," depending on how specific their vision is at that moment. This gap between how humans express desire and how machines have been trained to listen has left digital fashion assistants struggling to capture the true intent behind a request.
A team of researchers from Italy and Austria has now bridged this gap by teaching a new system to understand free-form instructions. Instead of forcing users to fit their desires into a pre-set box, the researchers created a method where a computer can take a seed image of a garment and a natural language sentence, no matter how detailed or vague, and generate a matching item. They tested this by enriching three existing fashion datasets with thousands of new instructions, ranging from minimal cues like "gym shorts" to highly detailed descriptions specifying fabric, color, and fit. The result is a system called StyleFlow, which does not just guess what a matching item might look like, but actually listens to the specific words used to describe it. In their experiments, the researchers found that the more specific the language provided, the better the computer could align its creation with the user's vision, producing garments that were not only stylistically compatible with the original item but also faithful to the written description.
The researchers began by addressing a fundamental limitation in previous work: most systems were trained and tested using fixed templates, such as "a photo of [category]." This approach failed to reflect how people actually shop, where queries vary wildly in detail. To fix this, the team developed a two-step process to generate realistic, free-form instructions for their tests. First, they used a vision-language model to analyze images of clothing and write structured captions describing attributes like color, material, and cut. Then, they asked the same model to rewrite those captions into natural-sounding sentences at three different levels of detail. One level offered only the basic category, another added color and general style, and the third included specific details about the fit and fabric. Human annotators then reviewed these thousands of generated instructions to ensure they were clear, fluent, and accurately described the visual content. This created a controlled environment where the researchers could test how well a computer handles everything from a vague hint to a precise order.
To put these instructions to the test, the researchers built StyleFlow, a new type of generative model based on a technique called Rectified Flow Matching. Unlike older systems that often required complex, separate modules to handle images and text, StyleFlow integrates both inputs into a single, unified structure. It takes the image of the seed garment and the text instruction and processes them together to generate a new image. The researchers trained this model on three large datasets containing tens of thousands of compatible top-and-bottom pairs, ensuring that the items used for testing were never seen during training. This setup forced the system to learn the underlying principles of style and compatibility rather than simply memorizing specific pairs. When the system generated a new garment, it was evaluated not just on how realistic the image looked, but also on how well it matched the text instruction and how closely it resembled actual items found in a real clothing catalog.
The results showed a clear and consistent pattern: the specificity of the language directly influenced the quality of the result. When the system received high-detail instructions, it produced garments that were significantly more aligned with the intended design and the original seed item. For instance, on one of the datasets, moving from a low-detail prompt to a high-detail one improved the system's ability to match the generated item to real catalog products from about 38 percent to nearly 69 percent. The system also performed better on standard measures of image quality, with the generated items looking more realistic and less blurry when given more descriptive text. In contrast, when the instructions were vague or missing entirely, the system's performance dropped sharply, producing items that were less compatible and harder to distinguish from random noise. This confirmed that the system relies heavily on the text to guide its creative process, using the seed image primarily to maintain stylistic harmony.
Human evaluators further validated these findings by comparing the new system against several established models. In a blind study, participants rated the visual quality, style compatibility, and authenticity of the generated images. The new system consistently outperformed its competitors, producing items that looked more realistic and were more likely to be mistaken for genuine catalog photos. Participants judged the system's creations as being nearly as authentic as real photographs, with a "fake" detection rate of only about 28 percent, compared to much higher rates for older models. The system also excelled at following the user's intent; when asked for specific features like a "zippered pocket" or "elastic waistband," it successfully incorporated those details, whereas other models often ignored them or produced generic versions of the requested item.
The study also explored what happens when the system is deprived of its inputs. When the researchers removed the text instructions, the system's ability to generate a relevant item collapsed, proving that the language is not just a minor addition but a core driver of the generation process. Similarly, when the seed image was replaced with a blank canvas, the system struggled to maintain stylistic coherence, especially when the text instructions were vague. This demonstrated that the system needs both the visual context of the seed item and the specific guidance of the text to work effectively. The researchers noted that while the system works best with detailed instructions, it remains capable of producing plausible results even with minimal cues, suggesting it could be useful for exploring design ideas when a user is unsure of exactly what they want.
By reframing the problem from a rigid template-filling exercise to a flexible language-grounding task, this work moves the field of fashion generation closer to how people actually interact with technology. The researchers showed that machines can learn to interpret the full spectrum of human language, from the briefest hint to the most elaborate description, and translate it into a visual reality that respects the original style. While the current system focuses on matching tops and bottoms, the authors suggest that this approach could eventually be expanded to handle more complex outfit combinations and real-world user queries. The study concludes that the key to better digital fashion assistants lies not in building more complex architectures, but in teaching them to listen more carefully to the words people use to describe what they want to wear.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.