Leave No Stone Unturned: Uncovering Holistic Audio-Visual Intrinsic Coherence for Deepfake Detection
This paper introduces HAVIC, a deepfake detector that leverages holistic audio-visual intrinsic coherence learned from authentic videos to achieve superior generalization and performance on a new high-fidelity dataset (HiFi-AVDF), significantly outperforming existing state-of-the-art methods in cross-dataset scenarios.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are a detective trying to spot a fake video. In the past, deepfake detectors were like detectives who only looked at one clue: either the picture (visuals) or the sound (audio). They looked for tiny, specific glitches left behind by old, clumsy forgers—like a blurry edge on a face or a weird echo in a voice.
But today's forgers are using super-advanced AI (like the ones making movies or generating videos from text). They are so good that they don't leave those old, obvious glitches. They are like master magicians who make the picture and sound look perfect individually, but they might still be slightly out of sync or just "feel" wrong when you put them together.
This paper introduces a new detective named HAVIC (Holistic Audio-Visual Intrinsic Coherence). Instead of looking for specific glitches, HAVIC looks for natural harmony.
Here is how it works, broken down into simple concepts:
1. The "Natural Harmony" Detective
Think of a real human conversation like a perfectly orchestrated orchestra.
- The Visuals: The singer's lips move exactly when the note is sung.
- The Audio: The voice sounds like it's coming from the person's mouth, not a speaker in the corner.
- The Context: If the singer is talking about a beach, the background looks like a beach.
Old detectors were like a technician checking if a single instrument was out of tune. HAVIC is like a conductor listening to the entire orchestra. It asks: "Does the sound match the movement? Does the movement match the scene? Does the whole thing feel like a real, living moment?"
2. The Three Levels of "Feeling Real"
The authors say HAVIC checks for "coherence" (harmony) on three levels, like checking a painting for authenticity:
- Level 1: The Single Modality Check (The Texture)
- Analogy: Looking at a painting to see if the brushstrokes look like real paint or if the texture is weird.
- What HAVIC does: It checks if the video alone looks natural (no weird face deformities) and if the audio alone sounds natural (no robotic distortions).
- Level 2: The Micro-Check (The Lip-Sync)
- Analogy: Watching a puppet show. If the puppet's mouth opens for the word "Hello," but the sound comes out a split second late, you know it's fake.
- What HAVIC does: It checks the tiny, millisecond timing between a sound (like a "p" or "b" sound) and the exact shape of the lips.
- Level 3: The Macro-Check (The Story)
- Analogy: Imagine a video of a person screaming in terror, but the audio is a cheerful lullaby. The story doesn't match.
- What HAVIC does: It checks the big picture. If the audio says "I'm at the beach," but the video shows a snowy mountain, HAVIC knows something is wrong, even if the face and voice look perfect.
3. How HAVIC Learned to Be a Detective
You can't teach a detective to spot fakes just by showing them fakes, because forgers keep changing their tricks. Instead, HAVIC was trained on thousands of real, authentic videos.
- The Training: Imagine showing a student millions of real videos and saying, "Learn what 'real' feels like." HAVIC learned the "laws of physics" for how sound and vision naturally interact.
- The Test: When it sees a new video, it compares it against those "laws of physics." If the video breaks the rules of natural harmony (even slightly), it flags it as a fake.
4. The New "Exam" (The Dataset)
The authors realized that old tests for deepfakes were too easy. They were like giving a math genius a test with only addition problems. The new forgers can easily pass those.
So, the team created a new, super-hard exam called HiFi-AVDF.
- What is it? A collection of videos made by the world's most advanced commercial AI tools (like Sora, Kling, and Veo).
- Why is it special? These videos are so realistic that even humans struggle to tell them apart from reality. It's the ultimate stress test for any detector.
5. The Results
When they put HAVIC to the test:
- Old Detectors: They stumbled. They got confused by the new, high-quality fakes.
- HAVIC: It aced the test. It caught the fakes that others missed by a huge margin (improving detection by nearly 10% on the hardest tests).
The Bottom Line
This paper argues that to catch the next generation of AI fakes, we can't just look for "glitches." We have to understand how the world naturally works.
HAVIC is a detective that doesn't just look for broken pieces; it listens for the song of reality. If the song is even a tiny bit off-key, it knows it's a fake. This makes it much harder for bad actors to fool it, no matter how advanced their AI tools become.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.