MADBench: A Benchmark for Modality-Aware Audio Deepfake Detection
This paper introduces MADBench, the first benchmark that distinguishes between speech and environmental audio to evaluate deepfake detection, revealing that existing models fail on both components and that manipulating background audio asymmetrically degrades speech detection performance.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are watching a video of a famous politician giving a speech. The video looks perfect: the lighting is right, the camera is steady, and the politician's face moves naturally. But something feels "off." The voice sounds slightly robotic, and the background noise—the rustling leaves, the distant traffic, the hum of the air conditioner—sounds like it was recorded in a different room. This is the world of deepfakes. For a long time, scientists focused on catching the "face swap" tricks, where someone's face is digitally pasted onto another person's body. But now, the game has changed. The new trick isn't changing the face; it's changing the sound.
In this new era, bad actors can keep a real video but swap out the voice and the background noise with computer-generated audio. This creates a "fake" that looks 100% real but says something completely different. The big question for scientists is: How do we catch these audio tricks? The old way of looking for fakes often treated all sound as one big blob, assuming that if the voice was fake, the whole audio was fake. But this paper argues that sound is actually made of two very different ingredients: the speech (what is being said) and the environment (the background noise). Just like a cake needs both flour and eggs, a video needs both speech and background noise to sound real. If you mess with one, it might leave a different kind of "crumb" than if you mess with the other.
This is where MADBench comes in. Think of MADBench as a giant, super-organized "spot the fake" training gym for computers. The researchers built a massive dataset of videos where they kept the original video footage exactly the same but swapped out the audio in four different ways: they kept everything real, they faked only the voice, they faked only the background noise, or they faked both. They even made some of the fake background noises match the scene (like birds chirping in a park) and some that didn't (like traffic noise in a quiet library) to see if that confused the computers.
When they tested the current "champions" of fake detection against this new gym, the results were surprising. The computers that were previously trained to spot fakes mostly failed; they couldn't tell the difference between real and fake audio at all. However, a different type of computer model—one that had learned to understand both pictures and sounds together from a huge amount of data—did much better. But here is the twist: these smart models were much better at catching the fake background noise than the fake voice. It turns out that the "crumbs" left by a fake voice are much harder to find than the crumbs left by a fake background.
The paper also discovered something tricky: if you fake the background noise, it actually makes it harder for the computer to spot the fake voice. It's like if someone puts a loud, fake siren in the background; it distracts the detective so much they miss the fake voice in the foreground. The researchers found that while these smart models could tell the difference between a real scene and a fake one (like knowing a siren doesn't belong in a library), they still struggled to pinpoint exactly which part of the audio was the lie.
In short, the paper suggests that our current tools are not ready for this new kind of audio trickery. We can't just look at the sound as one big thing anymore; we have to look at the voice and the background separately. The study shows that while we have some smart models that can spot when the background doesn't match the video, we still have a long way to go before we can reliably catch a fake voice, especially when it's hiding behind a fake background. The authors conclude that to build better detectors for the future, we need to stop treating audio as a single stream and start treating the voice and the environment as two separate suspects to be investigated.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.