AGIDefect-4K: A Richly Annotated Dataset for AI-Generated Image Defect Detection, Localization and Explanation
This paper introduces AGIDefect-4K, a comprehensive dataset of 4,000 images from 15 generative models featuring hierarchical annotations for defect detection, localization, and explanation, alongside AGIDA, a baseline framework leveraging Multimodal Large Language Models to address the underexplored challenge of AI-generated image defect diagnosis.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
In the last few years, computers have learned to paint pictures that look startlingly real. By reading a simple sentence, a machine can generate an image of a sunset over a mountain range or a portrait of a person in a crowded street. These images are often so convincing that they blur the line between what was captured by a camera and what was invented by code. Yet, despite this incredible progress, these machines are not perfect. They frequently make subtle mistakes that a human eye might miss at first glance but that betray the image as artificial. A hand might have too many fingers, a floating object might lack a shadow, or a face might twist into an impossible shape. These errors are not just minor glitches; they are cracks in the foundation of trust. If we cannot reliably spot these flaws, we cannot fully trust the images we see online, nor can we help the machines learn to do better.
A team of researchers at Xidian University in China has set out to map these invisible cracks. They realized that while we have many ways to judge how "pretty" an image looks, we lack a systematic way to find exactly where and why an image is broken. To solve this, they created a new collection of 4,000 images, a resource they call AGIDefect-4K. This dataset is not just a gallery of pictures; it is a detailed map of errors. The researchers gathered images from fifteen different artificial intelligence systems, ranging from open-source tools available to anyone to powerful, closed systems used by major companies. They then enlisted a team of human experts to examine every single image. These experts did not just say whether an image was good or bad. They performed a three-step inspection: first, they decided if a defect existed at all; second, they drew a precise outline around the specific area where the error occurred; and third, they wrote a detailed description explaining what was wrong and how much it hurt the overall quality of the picture.
The results of this meticulous work revealed a surprising truth. Even the most advanced machines, the ones that produce the most stunning visuals, still fail in predictable ways. The researchers found that defect rates varied between the different systems, with some producing errors in nearly 15 percent of their images, while others were slightly more reliable. The errors were not random noise; they fell into clear categories. Some were structural, such as limbs that did not connect properly or objects that defied the laws of physics by floating without support. Others were logical, where the combination of elements made no sense to human experience. What made this dataset unique was the depth of the information attached to each image. For every flaw, the researchers recorded not just a label, but a pixel-perfect mask showing the exact location of the problem and a paragraph of text describing the nature of the mistake. This level of detail allows for a much deeper understanding than simply counting how many images are "bad."
Using this rich collection of data, the researchers built a new tool called AGIDA, an assistant designed to look at an image and perform the same three-step inspection that the human experts did. They trained this tool to detect the presence of a defect, point to its location, explain what is wrong, and give the image a quality score. When they tested this tool against the best existing artificial intelligence systems, the results were telling. The most powerful general-purpose models, which can write poetry and answer complex questions, struggled significantly when asked to find these specific visual errors. They often missed obvious mistakes, such as a bird with only one leg, or they focused on the wrong details, describing a mouse's shape when the real problem was a fused finger on a human hand. In contrast, the new tool, trained specifically on the detailed human annotations, learned to spot these errors with much greater accuracy. It could identify the fused fingers and the missing leg, drawing a mask around the error that matched the human experts' work almost perfectly.
The study also showed that understanding these defects is not just about finding errors; it is about improving how we judge quality. The researchers found that the presence of a defect, and how severe it was, directly influenced the overall score an image received. When the tool was taught to pay attention to the specific defects it found, its ability to judge the overall quality of an image improved. This suggests that the path to better artificial intelligence images lies in teaching the machines to see their own mistakes. By providing a clear, detailed map of where things go wrong, the researchers have given the field a new standard for evaluation. They have shown that while the machines are getting better at creating art, they still need a human-like ability to notice when something is slightly, but critically, out of place. This new dataset and the tool built from it offer a way forward, turning the vague feeling that an image looks "off" into a concrete, measurable fact that can be fixed.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.