← Latest papers
🤖 AI

A Comprehensive Dataset for Human vs. AI Generated Image Detection

This paper introduces MS COCOAI, a comprehensive dataset comprising 96,000 real and synthetic images generated by five leading AI models, designed to advance research in detecting and identifying the source of AI-generated media.

Original authors: Rajarshi Roy, Ashhar Aziz, Shashwat Bajpai, Nasrin Imanpour, Gurpreet Singh, Shwetangshu Biswas, Kapil Wanaskar, Parth Patwa, Subhankar Ghosh, Shreyas Dixit, Nilesh Ranjan Pal, Vipula Rawte, Ritvik Ga
Published 2026-05-26
📖 4 min read☕ Coffee break read

Original authors: Rajarshi Roy, Ashhar Aziz, Shashwat Bajpai, Nasrin Imanpour, Gurpreet Singh, Shwetangshu Biswas, Kapil Wanaskar, Parth Patwa, Subhankar Ghosh, Shreyas Dixit, Nilesh Ranjan Pal, Vipula Rawte, Ritvik Garimella, Amitava Das, Amit Sheth, Gaytri Jena, Vasu Sharma, Aishwarya Naresh Reganti, Vinija Jain, Aman Chadha

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you walk into a massive art gallery. Half the paintings are real photographs taken by humans, and the other half are masterpieces created by a team of very talented, but invisible, robots. The problem? The robots are getting so good that you can't tell the difference just by looking.

This paper is about building a training camp and a test to help us spot the robot-made art. Here is the breakdown in simple terms:

1. The Problem: The "Uncanny Valley" of Pictures

We have powerful AI tools (like Stable Diffusion, DALL-E 3, and MidJourney) that can create images from scratch. They are amazing, but they are also being used to spread lies or fake news. Because these images look so real, our eyes (and even old computer programs) are getting fooled. We need a better way to catch them.

2. The Solution: The "MS COCOAI" Dataset

The authors built a giant library of 96,000 pictures to train computers to be better detectives. Think of this library as a controlled science experiment.

  • The "Real" Group: They took 16,000 real photos from a famous database called MS COCO. These are genuine photos with human-written descriptions (captions).
  • The "Fake" Group: They took those exact same human-written descriptions and fed them into five different AI robots (Stable Diffusion 2.1, SDXL, SD 3, DALL-E 3, and MidJourney v6).
  • The Magic Trick: Because the AI and the human photographer used the same description, the resulting pictures are "twins" in terms of content.
    • Analogy: Imagine asking a human chef and a robot chef to make a "chocolate cake with strawberries." If the robot makes a cake that looks slightly different, we know it's the robot's fault, not because the robot was asked to make a "spicy pizza" instead. This helps researchers separate "content bias" from "AI artifacts."

3. The Stress Test: "The Gymnastics Routine"

To make sure the detectors are tough, the authors didn't just give them clean pictures. They also applied four "stunts" to the AI images, like a gymnast adding difficulty to a routine:

  1. Flipping: Mirroring the image horizontally.
  2. Dimming: Making the image darker.
  3. Static: Adding TV-like "snow" (noise) to the picture.
  4. Compression: Squeezing the file size (like saving a photo on a phone) to blur it slightly.

This ensures that if a detector works, it works even when the image has been messed with.

4. The Two Challenges (The Tests)

The paper sets up two specific games for computers to play using this dataset:

  • Game A (The "Real or Robot?" Test): Look at a picture and say, "Is this human-made or AI-made?"
    • Result: A basic computer model got about 80% right. This is like a human guessing correctly 4 out of 5 times. It's doable, but not perfect.
  • Game B (The "Who Made It?" Test): Look at an AI picture and guess which of the five robots made it.
    • Result: The computer only got about 45% right. This is barely better than random guessing (like flipping a coin).
    • Why is this hard? It's like trying to tell the difference between five different professional forgers who are all trying to copy the same style. It is much harder to identify the specific artist than to just spot that the art is fake.

5. How They Tried to Solve It (The Baseline)

To get a starting point, the authors didn't use fancy new AI. They used a standard camera lens (a ResNet-50 model) but looked at the pictures in a different way: Frequency Domain.

  • Analogy: Imagine looking at a painting not by its colors, but by the "vibrations" or the sound it would make if it were a musical instrument. AI images often have a weird "hum" or a specific pattern in their vibrations that human photos don't have. They taught a computer to listen for that hum.

6. The Bottom Line

The paper concludes that while we are getting better at spotting that an image is fake (Game A), we are still terrible at figuring out which specific AI tool made it (Game B).

They have released this massive library of 96,000 paired images (real vs. AI) to the public so other researchers can try to build better detectors. Their goal is to help society trust what it sees online again.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →