PaletteGAN: Palette-Guided Controllable Colorization for Fashion Sketches
This paper proposes PaletteGAN, a palette-guided GAN framework that enhances structural preservation and color controllability in fashion sketch colorization by integrating brightness-preserving data augmentation, a perceptually optimized six-color palette construction, and a specialized generator-discriminator architecture, achieving superior performance in visual realism and design consistency on the Cleaned Maryland Dataset.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
In the world of fashion design, a sketch is merely a skeleton. It captures the shape of a garment, the drape of a sleeve, or the curve of a collar, but it lacks the life that color provides. For decades, turning these black-and-white drawings into full-color visions required the slow, deliberate hand of a human artist, who would test countless combinations to find the perfect harmony. Today, computers are learning to do this work, but they face a unique challenge: unlike a simple photograph, a fashion sketch is a complex map of folds, seams, and decorative lines that must remain perfectly intact while the computer invents the colors. If the machine gets the color right but blurs the lines, the result is useless to a designer. The goal is to create a system that can take a designer's specific color choices and apply them to a sketch with the precision of a human hand, preserving every structural detail while exploring new visual possibilities.
Researchers at Wuhan Textile University have developed a new tool called PaletteGAN to solve this exact problem. Instead of letting a computer guess what colors might look good, this system asks the designer to provide a specific set of six colors first. Think of this set as a small, curated box of paints. The computer then takes a fashion sketch and this specific box of colors and works to paint the garment, ensuring the final image looks realistic and that the colors stay exactly where the designer intended. The team found that by combining a specific way of drawing the sketch's lines with a smart method for choosing those six colors, they could generate images that are far more accurate and stable than previous attempts.
The process begins with the data itself. The researchers started with a large collection of real fashion photographs. To teach the computer how to handle different lighting and shades, they created thousands of variations of these photos by subtly shifting the colors while keeping the brightness exactly the same. From each photo, they extracted a sketch, but not just any sketch. They tested four different ways to draw the lines of the garment. One method simply traced the edges, while another tried to keep the internal folds and seams visible. The researchers discovered that the most successful approach was a hybrid technique that kept the strong outer outline of the clothing while also preserving the delicate, weak lines inside that show how the fabric folds. This specific type of sketch, which they named DC-HED-Contour, provided the clearest map for the computer to follow.
Next came the color selection. A common way to pick colors from an image is to group similar shades together, but this often misses small, important details like a tiny button or a specific trim. The team developed a new strategy to ensure the six colors in the palette were not just the most common ones, but also included distinct, contrasting shades that gave the garment its character. They used a mathematical rule to ensure that once a main color was chosen, the next color selected had to be as different from it as possible in human perception. This guaranteed that the final palette would always include a dominant color along with enough variety to capture the secondary and accent details of the design.
The core of their invention is a neural network, a type of computer program designed to learn from examples. This network has two main parts working in tandem. The first part, the generator, takes the sketch and the six-color palette and tries to paint the picture. The second part, the discriminator, acts as a strict critic. It does not just look at the final image to see if it looks real; it also checks if the lines match the original sketch and if the colors match the provided palette. By constantly comparing its own work against these strict rules, the generator learns to produce images that are not only colorful but also structurally perfect. The researchers tested this system against several other existing methods and found that their approach produced the most realistic results. When they measured how close the generated images were to real photographs, their method scored significantly better than the competition.
Perhaps the most important finding was not just about the average quality of the images, but about the consistency. Other systems sometimes produced a few perfect images but many others that were blurry or had strange color patches. The new system, however, delivered a steady, high-quality result across almost every single test case. It successfully preserved the complex folds of a dress and the sharp lines of a collar while applying the designer's chosen colors without smearing or bleeding. While the system is not perfect—it can sometimes struggle with extremely complex textures or very subtle gradients—it represents a significant step forward. It proves that by giving the computer a clear structural map and a specific, well-chosen set of colors, we can automate the creative process of fashion design without losing the human touch that makes a sketch come to life.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.