← Latest papers
💻 computer science

Skyra: AI-Generated Video Detection via Grounded Artifact Reasoning

This paper introduces Skyra, a specialized multimodal large language model that detects and explains AI-generated videos by identifying human-perceivable visual artifacts, supported by the new ViF-CoT-4K dataset, a two-stage training strategy, and the comprehensive ViF-Bench evaluation.

Original authors: Yifei Li, Wenzhao Zheng, Yanran Zhang, Runze Sun, Yu Zheng, Lei Chen, Jie Zhou, Jiwen Lu

Published 2026-05-18
📖 4 min read☕ Coffee break read

Original authors: Yifei Li, Wenzhao Zheng, Yanran Zhang, Runze Sun, Yu Zheng, Lei Chen, Jie Zhou, Jiwen Lu

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine the internet is suddenly flooded with videos that look so real, you can't tell if they were filmed with a camera or dreamed up by a computer. This is the problem Skyra was built to solve.

Think of Skyra not as a simple "yes/no" detector, but as a super-powered video detective who doesn't just guess if a video is fake, but actually points a finger at the specific "glitch" that gives it away and explains why it's a glitch.

Here is how the paper breaks down, using some everyday analogies:

1. The Problem: The "Uncanny Valley" is Getting Smoother

AI video generators (like Sora, Kling, or Wan) have gotten incredibly good. They can make videos that look perfect at a glance.

  • The Old Way: Previous detectors were like security guards with a checklist. They looked for "bad pixels" or "weird lighting." But as AI gets better, those bad pixels disappear. The guards started getting confused, often saying "Fake!" to real videos or "Real!" to fakes.
  • The MLLM Problem: Even the smartest general AI chatbots (the "geniuses" of the AI world) failed this test. When asked to spot a fake video, they would write long, confident essays saying, "The lighting looks great, the person is smiling, so this is real!" They missed the subtle, physics-breaking errors because they weren't trained to look for them.

2. The Solution: Skyra, the "Forensic Art Critic"

Skyra is a specialized AI model designed to do one thing: find the "artifacts."

  • What is an artifact? Imagine you are looking at a painting. If the artist painted a hand with six fingers, or if a cup of coffee suddenly turned into a bird in mid-air, that's an artifact. It's a mistake that breaks the laws of physics or common sense.
  • How Skyra works: Instead of just saying "Fake," Skyra acts like a detective with a magnifying glass. It watches the video frame-by-frame and says:

    "At 1.5 seconds, the bartender's metal shaker suddenly melted like butter. That's impossible! That's a Shape Distortion. Also, at 3.0 seconds, a glass appeared out of thin air. That's an Abnormal Object Appearance."

3. The Training: Teaching the Detective with a "4,000-Video Case File"

You can't teach a detective to spot fakes just by asking them to guess. You need real cases.

  • The Dataset (ViF-CoT-4K): The researchers built a massive library of 4,000 videos. Crucially, they didn't just label them "Fake." They had human experts watch them and write detailed reports on exactly what was wrong, down to the specific second and the exact spot on the screen.
  • The "Chain of Thought" (CoT): They taught the AI to think out loud. Instead of jumping to a conclusion, the AI is forced to say, "I see this, then I see that, therefore this is fake." This mimics how a human expert reasons.

4. The Two-Step Training: "Cold Start" and "Reinforcement"

The paper describes a clever two-step training process:

  • Step 1: The Cold Start (SFT): They first taught the AI the basics using the human-written reports. It's like giving a new detective a textbook of "Common Mistakes in Fake Videos" so they know what to look for.
  • Step 2: Reinforcement Learning (RL): This is the "practice" phase. The AI is tested, and if it misses a clue or gets the reasoning wrong, it gets a "penalty." If it finds a subtle, physics-breaking error that even humans might miss, it gets a "reward." This pushes the AI to become a master detective who can spot the tiniest, most subtle errors.

5. The Results: Beating the Competition

The researchers tested Skyra against a "ViF-Bench," a tough test suite containing videos from over 10 of the world's most advanced AI video generators.

  • The Outcome: Skyra didn't just win; it dominated. While other detectors (and even the smartest general AI models) struggled, often performing no better than random guessing, Skyra achieved very high accuracy.
  • The "Why": The paper shows that Skyra succeeds because it focuses on intrinsic violations (things that break physics, like a person walking through a wall or an object changing color instantly) rather than superficial things like "does this look grainy?"

Summary

In short, Skyra is a specialized AI detective trained on a massive, human-annotated library of "AI mistakes." It doesn't just guess if a video is fake; it watches the video, spots the specific moment where physics breaks (like a melting spoon or a disappearing person), and writes a report explaining exactly why it knows the video is a fabrication. It is designed to be the "truth-teller" in a world where seeing is no longer believing.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →