Artifact-Bench: Evaluating MLLMs on Detecting and Assessing the Artifacts of AI-Generated Videos
This paper introduces Artifact-Bench, a comprehensive benchmark featuring a three-level artifact taxonomy and three diagnostic tasks to evaluate Multimodal Large Language Models (MLLMs) on AI-generated video artifact detection, revealing significant limitations in their perceptual accuracy and alignment with human preferences.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are a detective trying to spot a fake painting. You know the artist is talented, but sometimes they make tiny mistakes: a hand has six fingers, a shadow falls the wrong way, or a clock runs backward.
This paper, Artifact-Bench, is like a new, super-tough test for "AI Detective" robots (called Multimodal Large Language Models, or MLLMs). The researchers wanted to see if these smart AI robots are actually good at spotting the tiny, weird mistakes (artifacts) that happen when computers make fake videos.
Here is the breakdown of what they did and what they found, using simple analogies:
1. The Problem: The "Uncanny Valley" of Video
AI video generators have gotten amazing. They can make videos that look almost real. But they aren't perfect. They still leave behind "glitches" or "artifacts."
- The Glitches: Sometimes a person's face melts, a car drives through a wall, or a person's arm flickers in and out of existence.
- The Question: Can current AI models (the "detectives") actually see these glitches and explain why the video looks fake, or are they just guessing?
2. The Solution: A New "Driving Test" for AI
The researchers built a special test called Artifact-Bench. Think of it as a three-part driving exam for AI detectives:
- Level 1: The "Real or Fake?" Quiz (Classification)
- The Task: Show the AI one video. Ask: "Is this real or AI-made?"
- The Catch: The AI has to guess based on visual glitches, not just the story. (e.g., If a video shows a dog flying, that's a story clue. If the dog's fur looks like static noise, that's a glitch clue).
- Level 2: The "Which is Better?" Showdown (Comparison)
- The Task: Show the AI two fake videos. Ask: "Which one looks more realistic?"
- The Catch: Both are fake, but one has fewer glitches. The AI has to spot the subtle differences.
- Level 3: The "Forensic Report" (Diagnosis)
- The Task: Show the AI a fake video and ask: "What exactly is wrong with this?"
- The Catch: The AI must pick the specific mistake from a list of 30 options (like "The water flowed uphill" or "The person's face melted"). This is the hardest part because it requires explaining the why, not just saying "it's fake."
3. The Map: A New "Glitch Dictionary"
Before the test, the researchers created a massive "Glitch Dictionary" (a taxonomy). They organized all possible mistakes into three levels:
- Surface Level: Things that look weird immediately (bad colors, blurry textures).
- Structure Level: Things that don't make sense physically (a chair with no legs, a person merging into a wall).
- Time & Logic Level: Things that break the laws of physics or time (a cup un-breaking itself, a person walking backward while moving forward).
4. The Results: The Detectives Failed the Exam
The researchers tested 19 of the smartest AI models available today. Here is what happened:
- The "Coin Flip" Problem: On the easiest tasks, many AI models did no better than if they had just flipped a coin. They couldn't reliably tell real videos from fake ones.
- The "Diagnosis" Disaster: On the hardest task (explaining why a video is fake), the models were terrible. Their accuracy dropped to near zero (often below 10%). It's like a student who can guess "True or False" but fails completely when asked to write an essay explaining the answer.
- The "Thinking" Trap: The researchers tried using "thinking" models (AI that talks through its reasoning) and bigger models. Surprisingly, bigger or "thinking" models didn't necessarily get better. Sometimes, they got worse. It turns out, just having a bigger brain doesn't help if you don't have the right "eyes" to see tiny details.
- The Human Gap: Human experts were great at the test. The AI models were not. The AI's judgment often didn't match what humans thought was realistic.
5. The Conclusion: The AI is "Blind" to the Details
The paper concludes that while AI is great at understanding the story of a video, it is currently bad at spotting the tiny, physical glitches that make a video look fake.
- The Metaphor: Imagine an AI that can describe a movie plot perfectly but is blind to the fact that the actor's hand is holding a fork instead of a sword.
- The Takeaway: We cannot yet trust these AI models to be the "judges" of video quality. If we use them to grade other AI videos, they might give a passing grade to a video full of weird glitches because they aren't looking closely enough.
In short: AI video generators are getting better, but the AI models we use to check their work are still struggling to see the "glitches in the matrix." We need smarter detectives before we can trust them to police the future of fake videos.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.