← Latest papers
💻 computer science

VideoASMR-Bench: Can AI-Generated ASMR Videos Fool VLMs and Humans?

The paper introduces VideoASMR-Bench, a novel benchmark utilizing fine-grained ASMR content to reveal that while state-of-the-art video generation models can create synthetic videos that fool current video understanding models, humans remain significantly better at distinguishing these AI-generated videos from real ones.

Original authors: Jiaqi Wang, Weijia Wu, Yi Zhan, Rui Zhao, Ming Hu, James Cheng, Wei Liu, Philip Torr, Kevin Qinghong Lin

Published 2026-05-08
📖 4 min read☕ Coffee break read

Original authors: Jiaqi Wang, Weijia Wu, Yi Zhan, Rui Zhao, Ming Hu, James Cheng, Wei Liu, Philip Torr, Kevin Qinghong Lin

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are holding a very sensitive "sensory test" in your hands. Most video tests today are like checking if a painting looks like a landscape from far away—they ask, "Is there a tree? Is there a sky?" But this new paper, VideoASMR-Bench, asks a much harder question: "Does the sound of the rain actually match the movement of the leaves? Does the texture of the water feel real when you watch it?"

Here is the breakdown of what the researchers did, using simple analogies:

1. The "ASMR" Test: A Microscope for Reality

The researchers decided to test AI videos using ASMR (Autonomous Sensory Meridian Response). Think of ASMR videos as the "high-definition, slow-motion" of the video world. They feature close-ups of things like tapping on glass, cutting fruit, or whispering.

  • Why this matters: In normal videos, a small glitch might go unnoticed. In an ASMR video, if a knife doesn't slice a tomato perfectly or the sound of the crunch is slightly off, it breaks the "spell" immediately. It's like trying to spot a fake diamond under a magnifying glass; the flaws are tiny but obvious if you know what to look for.

2. The Game: The Forger vs. The Detective

The paper sets up a giant game of "Cat and Mouse" between two types of AI:

  • The Forgers (Video Generation Models): These are the AI artists trying to create fake ASMR videos that look and sound so real they can trick anyone.
  • The Detectives (Video Understanding Models): These are the AI "police" trying to watch the videos and say, "This is real!" or "This is fake!"

The researchers created a massive library of 1,500 real ASMR videos (scraped from social media) and used the "Forgers" to create 2,235 fake versions of them. Then, they let the "Detectives" try to spot the fakes.

3. The Shocking Results

The results were surprising, like finding out that a master forger can fool a high-tech security camera, but a regular human can still spot the fake.

  • The AI Detectives are failing: Even the smartest AI detectives (like the latest versions of Gemini and GPT) are terrible at spotting these fake ASMR videos. They often guess randomly or get tricked easily. In fact, they are much worse than humans at this specific task.
  • The AI Forgers are getting scary good: The AI artists (like Veo3 and Sora2) are creating videos that are so convincing that the AI detectives can't tell the difference. The videos look and sound perfect to the machines.
  • Humans still win (for now): While the AI detectives are confused, regular humans can still usually tell the fake videos apart. Humans are better at noticing the tiny, subtle "uncanny valley" feelings that the AI misses.

4. The "Cheat Codes" the AI Used (and why they failed)

The researchers found out why the AI detectives were failing:

  • The Watermark Shortcut: Some AI models were just looking for a tiny "Sora" watermark on the video. If they saw it, they said "Fake!" If they didn't, they said "Real!" They weren't actually watching the video; they were just looking for a label. When the researchers removed the watermarks, the AI detectives got confused and failed.
  • Ignoring the Sound: The AI detectives mostly ignored the audio. But in ASMR, the sound is half the battle. When the researchers forced the AI to listen to the audio, they got slightly better at spotting fakes, but still not great.
  • The "Real" Bias: The AI detectives had a weird habit of assuming everything was real. They were so afraid of calling a real video "fake" that they ended up calling almost everything "real," even the obvious fakes.

5. The Big Picture

The paper concludes that while AI is getting incredibly good at making realistic sensory videos, the AI tools we use to check if those videos are real are lagging behind.

Think of it like this: The AI artists have learned to paint a perfect sunset that even the sun itself would be jealous of, but the AI security guards checking the paintings are still using a magnifying glass from the 1990s. They can't see the tiny brushstrokes that give the painting away.

The Bottom Line:
We have a new benchmark (VideoASMR-Bench) that acts as a stress test. It shows that right now, AI can make videos that fool other AIs, but humans are still the best judges of what feels truly real. The paper doesn't promise these videos will be used for anything specific yet; it just says, "Look how good the fakers are getting, and look how bad our detectors are at catching them."

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →