VTONQA: A Multi-Dimensional Quality Assessment Dataset for Virtual Try-on
This paper introduces VTONQA, the first multi-dimensional quality assessment dataset for virtual try-on containing 8,132 images and 24,396 human ratings across three dimensions, which is used to benchmark existing VTON models and image quality metrics to reveal their limitations and guide future improvements.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
In the bustling world of online shopping, a familiar frustration persists: the gap between a garment on a screen and how it might look on a real person. To bridge this gap, technology has developed virtual try-on systems, which act as digital fitting rooms. These tools take a photograph of a person and a separate image of a piece of clothing, then attempt to merge them into a single, realistic picture showing the person wearing that outfit. While these systems have become common in e-commerce and digital fashion, they are not perfect. Often, the results look strange; clothes might appear stretched or torn, body shapes can warp unnaturally, or the fabric might not drape correctly. For these tools to become truly useful, developers need a way to measure exactly how good or bad a result is, moving beyond simple pixel comparisons to understand what a human eye actually perceives as a successful try-on.
A team of researchers at Shanghai Jiao Tong University has addressed this need by creating a new resource called VTONQA, the first large-scale dataset designed specifically to judge the quality of these virtual try-on images. Instead of relying on a single overall score, the researchers broke down the evaluation into three distinct dimensions: how well the clothing fits the body, how naturally the body shape and pose remain consistent, and the overall aesthetic quality of the final image. To build this, they gathered 8,132 images generated by 11 different virtual try-on models, ranging from classic computer vision techniques to modern artificial intelligence systems. These images were created by pairing 184 different people—representing a diverse mix of ages, genders, ethnicities, and body types, including children, the elderly, and pregnant individuals—with 80 different garments across eight categories, from t-shirts and sweaters to dresses and trousers.
Once the images were generated, the researchers did not leave the judging to machines. They recruited 40 human volunteers to act as the ultimate arbiters of quality. These volunteers, who had backgrounds in undergraduate and graduate education, were trained to look at each image and assign a score from one to five for each of the three dimensions. They evaluated whether the shirt looked like it was actually being worn, if the person's arms and legs still looked like human limbs, and if the final picture felt pleasing to look at. This process resulted in nearly 24,400 individual scores, creating a rich, detailed map of how humans perceive success and failure in virtual try-on. The researchers then used this human data to test how well existing computer programs could predict these scores on their own.
The findings revealed a significant gap between what computers currently measure and what humans actually see. Traditional methods used to judge image quality, which often focus on how closely pixels match between two images or how much noise is present, failed to align with human opinion. These standard tools could not detect the subtle distortions that make a virtual outfit look fake, such as a sleeve twisting unnaturally or a waistline disappearing. Even more advanced artificial intelligence models, which had been trained on general images, struggled when faced with the specific challenges of virtual try-on. The study showed that while some modern systems performed better than others, none were perfect. The closed-source commercial systems tested generally outperformed the open-source research models, but even the best open-source models showed inconsistency, often doing well with upper-body clothing like shirts but struggling significantly with lower-body items like trousers and skirts.
One of the most revealing aspects of the study was how different types of clothing and different body types affected the results. The data suggested that current technology handles the contours of the upper body and full-body dresses more reliably than it handles the complex shapes of the lower body. Similarly, while the systems worked reasonably well across various body shapes and ethnicities, the performance varied enough to show that no single model is yet a universal solution. The researchers found that the overall quality of an image was often more dependent on whether the body looked natural than on whether the clothing fit perfectly, suggesting that preserving the human form is the most critical factor for a convincing result. By providing this detailed, human-verified dataset, the researchers have given the field a new standard for measurement. This resource allows developers to see exactly where their models fail, moving the technology from a trial-and-error phase toward a future where digital fitting rooms can offer a truly realistic and reliable experience for shoppers.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.