← Latest papers
🤖 AI

Preserving Forgery Artifacts: AI-Generated Video Detection at Native Scale

This paper addresses the limitations of existing AI-generated video detection methods by introducing a large-scale dataset of 140K+ videos and a novel Qwen2.5-VL-based framework that operates natively at variable resolutions to preserve high-frequency forgery artifacts lost in conventional preprocessing.

Original authors: Zhengcen Li, Chenyang Jiang, Hang Zhao, Shiyang Zhou, Yunyang Mo, Feng Gao, Fan Yang, Qiben Shan, Shaocong Wu, Jingyong Su

Published 2026-04-07
📖 4 min read☕ Coffee break read

Original authors: Zhengcen Li, Chenyang Jiang, Hang Zhao, Shiyang Zhou, Yunyang Mo, Feng Gao, Fan Yang, Qiben Shan, Shaocong Wu, Jingyong Su

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

🎬 The Problem: The "Pixel Shredder"

Imagine you are a detective trying to spot a fake painting. The forgery has a tiny, almost invisible scratch on the canvas that only the real artist would know to look for.

Now, imagine the police force (current AI detectors) has a rule: "Before you look at the painting, you must shrink it down to the size of a postage stamp."

When you shrink a high-definition painting to a postage stamp, that tiny, crucial scratch gets blurred out or disappears entirely. You are left with just the general shape of the face, which looks perfect. The detective fails because the evidence was destroyed by the shrinking process.

This is exactly what current AI video detectors do. They take high-quality, realistic AI videos and force them into a small, fixed size (like 224x224 pixels) before analyzing them. In doing so, they throw away the "high-frequency artifacts"—the tiny, digital fingerprints left behind by AI generators.

🛠️ The Solution: The "Native-Scale" Detective

The authors of this paper say: "Stop shrinking the evidence!"

They built a new system that looks at the video at its original size and shape, no matter how big or long it is. They call this "Native Scale."

Think of it like this:

  • Old Way: You try to read a book by squinting at it through a keyhole. You miss the details.
  • New Way: You put on a pair of high-powered reading glasses and look at the book exactly as it was printed. You can see the texture of the paper and the tiny ink smudges.

🧩 How It Works: The "Lego" Analogy

The new system uses a special brain called Qwen2.5-ViT. Here is how it processes a video without ruining it:

  1. No Cutting or Stretching: Instead of forcing a wide video to fit a square box (which stretches the people and makes them look weird), the system respects the video's original shape.
  2. 3D Patching: Imagine the video is a giant 3D block of Lego bricks.
    • Old detectors chop the block into tiny, uniform squares, losing the 3D structure.
    • This new detector cuts the block into 3D chunks (width, height, and time) that match the video's natural flow. It keeps the "time" dimension intact, so it can see if a person's hand moves unnaturally from one second to the next.
  3. The Magic Lens: Because it sees the video in its original glory, it can spot:
    • Tiny Glitches: Like a flickering shadow or a weird texture on a wall that only exists in AI.
    • Time Travel Errors: Like a person blinking in a way that doesn't match the physics of the real world.

📚 The New Evidence Locker (The Dataset)

Detectives are only as good as their training. The authors realized that old training data was like a library of blurry, low-quality photos from 5 years ago. Modern AI videos are like 4K IMAX movies.

So, they built a massive new library called Magic Videos:

  • They collected 140,000+ videos from 15 different top-tier AI generators (both open-source and commercial).
  • They created a special "Test Drive" (the Magic Videos benchmark) using the absolute latest, most realistic AI generators to see if their detector could actually spot the fakes.

🏆 The Results: Why It Matters

When they tested their new "Native Scale" detective against the old ones:

  • The Old Detectives: Got confused. They missed the fakes because the evidence was too small to see.
  • The New Detective: Crushed it. By keeping the video at its full resolution, it found the subtle "digital fingerprints" that others missed.

The Big Takeaway:
To catch the next generation of AI fakes, we can't just use the old tools. We have to stop shrinking the evidence. We need to look at the video exactly as it is, in all its high-definition, variable-sized glory, to catch the tiny lies hidden in the pixels.

⚡ In a Nutshell

  • The Flaw: Current AI detectors shrink videos, accidentally deleting the clues needed to spot fakes.
  • The Fix: A new system that analyzes videos at their original size and original shape.
  • The Result: A much smarter detector that can spot even the most realistic AI videos because it never throws away the evidence.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →