← Latest papers
💻 computer science

FakeI2V-Bench: Benchmarking the Applicability of Image-level Deepfake Detectors for Deepfake Video Detection

The paper introduces FakeI2V-Bench, a comprehensive benchmark comprising nearly 100,000 videos generated by advanced models to evaluate deepfake detectors, revealing that image-level detectors can outperform video-level ones and proposing the IV-Bridge framework to significantly enhance their video detection capabilities.

Original authors: Pei Li, Sihan Chen, Delong Ran, Tianshuo Cong

Published 2026-08-05
📖 4 min read☕ Coffee break read

Original authors: Pei Li, Sihan Chen, Delong Ran, Tianshuo Cong

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are walking through a digital forest where trees can grow from thin air, rivers can change their course with a whisper, and faces can be swapped like masks on a mannequin. This is the world of Artificial Intelligence (AI) video generation. For a long time, these "deepfake" videos were easy to spot because they looked a bit like a bad cartoon. But now, the AI has grown up. It can create movies so real that your eyes can't tell the difference between a real person and a digital ghost. This is a big problem. If we can't tell what's real, we might believe fake news, trust fake evidence in court, or get tricked by scammers. To stop this, scientists build "detectors"—digital security guards trained to spot the tiny, invisible glitches that AI leaves behind. For years, these guards were trained to look at single, still pictures. But now, the bad guys are making moving movies. The big question is: Can a guard trained to spot a fake photo do a good job guarding a fake movie?

This is exactly the mystery a team of researchers set out to solve in their new paper, FakeI2V-Bench. They built a massive testing ground, a "gym" for AI detectors, filled with nearly 100,000 videos created by the latest and most powerful AI tools. They wanted to see if the old photo-detectors could be upgraded to catch video fakes, or if they needed entirely new guards.

Here is what they found:

The Photo Guards Struggle in the Movie Theater
First, the researchers tested the standard photo detectors. These are the guards who have only ever looked at still images. When they tried to watch a video, they got confused. It's like asking a person who has only ever judged a single brick to judge a whole moving castle. The video has motion, blurring, and weird flickering that a still photo doesn't have. When these photo detectors looked at individual frames of a video, their performance dropped significantly. Some of them, which were 98% accurate on photos, fell to around 55% on videos—basically guessing like a coin flip.

The Magic of the "Frame-to-Video" Bridge
But the story doesn't end with failure. The researchers realized that while a single frame might fool the photo detector, the whole video tells a different story. They tried a clever trick: instead of just looking at one frame, they made the detector watch the whole video and then combined all its opinions. They tested six different ways to combine these opinions, like taking the average, the highest score, or the middle score.

Surprisingly, this simple trick worked wonders. One photo detector, when it started "watching" the whole video and combining its thoughts, jumped from a 71% success rate to over 80%. In fact, this upgraded photo detector actually beat the best specialized video detector currently available, which only reached about 79%. This proved that photo detectors have huge potential, they just needed the right way to look at moving pictures.

The Super-Tool: IV-Bridge
To make these photo detectors even better, the team built a special toolkit called IV-Bridge. Think of this as a two-step training camp:

  1. Video-Frame Fine-Tuning (VFT): They took the photo detectors and gave them a crash course specifically on video frames. This helped them learn the "language" of moving images, like how light changes when a person turns their head.
  2. Multi-Mode Aggregation (MMA): They didn't just use one way to combine the video opinions. They used a smart computer program (a Random Forest) to listen to six different ways of combining the scores and decide which one was best for that specific video.

The results were incredible. After using IV-Bridge, 11 out of the 12 photo detectors they tested became stronger than the best video-only detectors. The star of the show, a detector named RINE-IV, achieved a success rate of 93.80%, smashing the previous record of 79.99%.

Why This Matters
The researchers also looked at how much "brain power" these detectors needed. The specialized video detectors were like heavy, expensive tanks—slow and requiring massive computers. The upgraded photo detectors, however, were like agile motorcycles. They were much faster and lighter, often running in just a few milliseconds, yet they were more accurate.

In short, this paper shows that we don't necessarily need to invent brand-new guards for the video world. We can take the excellent guards we already have for photos, give them a little video training, and teach them how to listen to the whole story. This makes catching deepfake videos faster, cheaper, and much more effective, offering a powerful new shield for our digital world.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →