Beyond Seeing Is Believing: On Crowdsourced Detection of Audiovisual Deepfakes
This paper evaluates crowdsourced detection of audiovisual deepfakes, finding that while aggregating multiple worker judgments can reliably screen for video authenticity, crowdsourcing struggles to consistently identify specific manipulation types and timestamps, particularly for complex audio-video cases.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are the editor of a massive newsroom. Suddenly, a flood of videos arrives, some real and some "deepfakes"—videos that look and sound real but have been secretly altered by AI. Your job is to sort them out. But you can't watch every single one alone, so you hire a team of 240 regular people (a "crowd") to help you spot the fakes.
This paper is the report card on how well that team performed. Here is what they found, explained simply:
The Setup: Two Different "Training Camps"
The researchers didn't just test the crowd once; they ran two separate experiments using two different sets of videos:
- The "Hard" Camp (AV-Deepfake1M): These videos were short, tricky, and very realistic.
- The "Easier" Camp (TMC): These videos were longer and slightly easier to spot.
In both camps, the crowd had to do three things for each video:
- Spot the Fake: Is this real or doctored?
- Name the Cheat: Did they mess with the audio (voice), the video (face), or both?
- Pinpoint the Time: Exactly when did the cheating start?
The Results: What the Crowd Got Right (and Wrong)
1. The "Safety Net" Effect (RQ1)
The crowd was excellent at not crying wolf. If a video was truly real, the workers almost never called it a fake. They were very careful not to accuse innocent videos.
- The Problem: They were terrible at catching the actual fakes. It was like a security guard who is very good at letting real people in, but misses almost all the thieves trying to sneak in.
- The Difference: The crowd caught about twice as many fakes in the "Easier" camp (TMC) compared to the "Hard" camp. This proves that the type of video matters a lot.
2. The "Herd Mentality" (RQ2)
When the researchers asked, "Did everyone agree?", the answer was "Not really."
- For any single video, the workers often disagreed with each other. One person might think it's real, while another thinks it's fake.
- However, when you take the majority vote (what most people said), the signal became much clearer. It's like asking a room full of people to guess the weight of a cow; one person might be way off, but the average of 10 people is usually pretty close.
- The Catch: Even with the majority vote, if the video was a really tricky fake, the crowd still missed it. You can't fix a blind spot just by asking more people.
3. The "Blind Spot" on Details (RQ3)
This was the hardest part. Even when the crowd did spot a fake, they were terrible at guessing how it was faked.
- The Mix-Up: If a video had both a fake voice and a fake face, the crowd usually just guessed "Fake Voice" or "Fake Face," rarely getting the "Both" answer right.
- The Silver Lining (Timestamps): While they couldn't agree on what was faked, they were surprisingly good at agreeing on when it happened. If they thought a video was fake, they could usually point to the same 5-second chunk where the weirdness started.
The Takeaway: A Two-Step Strategy
The paper suggests a practical way to use this "crowd" system, using a simple analogy:
Think of the crowd as a metal detector at an airport.
- Step 1 (Screening): The metal detector (the crowd) is great at scanning thousands of bags quickly. It rarely flags a bag that has nothing in it (low false alarms), but it might miss some small, hidden items (missed fakes).
- Step 2 (Expert Review): If the detector beeps, you don't ask the detector what the item is or where exactly it is inside the bag. You send that bag to a human expert (or a smart AI) to open it up and inspect it closely.
In short: The crowd is a great "first filter" to catch obvious fakes and flag suspicious videos for experts to look at. But don't ask them to be the final judge on how the video was faked, because they will likely get that wrong.
One Small Warning
The researchers noted that some workers might have started the test with their volume turned down. Since they couldn't hear the audio, they might have missed "audio-only" fakes. This is a reminder that even in a simple task, small details (like volume) can change the results.
Bottom Line: Crowdsourcing is a useful tool for spotting "something is wrong" in videos, but it needs help to figure out exactly "what is wrong."
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.