← Latest papers
⚡ electrical engineering

Benchmarking Single-Factor Physical Video-to-Audio Generation

This paper introduces FlatSounds, a benchmark that evaluates video-to-audio generation models on their physical reasoning capabilities through controlled counterfactual tests, revealing that while text captions improve semantic accuracy, they often degrade temporal alignment and that current models rely too heavily on text rather than learning physical processes directly from visual data.

Original authors: Tingle Li, Siddharth Gururani, Kevin J. Shih, Gantavya Bhatt, Sang-gil Lee, Zhifeng Kong, Arushi Goel, Gopala Anumanchipalli, Ming-Yu Liu

Published 2026-05-29
📖 4 min read☕ Coffee break read

Original authors: Tingle Li, Siddharth Gururani, Kevin J. Shih, Gantavya Bhatt, Sang-gil Lee, Zhifeng Kong, Arushi Goel, Gopala Anumanchipalli, Ming-Yu Liu

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are teaching a robot to listen to the world. You show it a video of someone hitting a glass jar with a spoon, and the robot makes a sound. If the sound is a nice "clink," we usually say the robot did a good job. But what if the robot is just guessing? What if it's not actually "seeing" the glass or the spoon, but just reading a label that says "glass jar" and playing a pre-recorded sound from its memory?

This paper, titled FlatSounds, asks a tough question: Do current AI video-to-audio models actually understand physics, or are they just good at guessing?

Here is the breakdown of their findings using simple analogies:

1. The Problem: The "Cheat Sheet" Robot

Current AI models are like students who are great at memorizing answers but terrible at understanding the math.

  • The Old Way: If you show a video of a metal spoon hitting a glass, the AI makes a "clink." If you show a wooden spoon hitting the same glass, the AI might make a duller "thud," but often it doesn't.
  • The Discovery: The authors found that these models rely heavily on text captions (the "cheat sheet"). If you tell the AI "a metal spoon hits a glass," it makes a metal sound. But if you take away the text and just show the video, the AI often fails to notice the difference between metal and wood. It's like a student who can only solve a problem if the teacher whispers the answer key; without the key, they don't know how the numbers work.

2. The New Test: "FlatSounds"

To test if the robots are actually smart or just memorizing, the researchers built a new test called FlatSounds. Think of this as a "physics lab" for AI.

They created two types of experiments:

  • The "What If?" Test (Counterfactuals): They took a video of a person tapping a jar full of sand and a video of a person tapping the same jar but full of water. They carefully slowed down or sped up the videos so the tapping happened at the exact same time.
    • The Test: If the AI is smart, the sound for the water jar should sound "heavier" or "duller" than the sand jar.
    • The Result: Most AIs failed. They sounded the same because they weren't looking at the liquid; they were just looking at the text label.
  • The "Pattern" Test: They showed videos where a piano player moves from low notes to high notes.
    • The Test: The sound should get higher in pitch.
    • The Result: The AI struggled to keep the pitch rising correctly just by watching the video.

3. The Big Surprise: The "Text vs. Timing" Trade-off

The paper found a strange, frustrating trade-off.

  • With Text: When the AI is allowed to read a description (e.g., "metal spoon"), the sound is semantically correct (it sounds like the right object) but out of sync. The "clink" happens a split second too late or too early. It's like a movie where the actor's mouth moves, but the sound effect is slightly delayed.
  • Without Text: When you force the AI to rely only on the video, the sound becomes perfectly timed (it clicks exactly when the spoon hits), but the sound itself is often wrong (it might sound like plastic instead of metal).

The Metaphor: Imagine a drummer.

  • Model A (With Text): Knows exactly what instrument to play (a snare drum), but hits the drum slightly off-beat.
  • Model B (No Text): Hits the drum perfectly on the beat, but sometimes plays a cymbal when they should be playing a snare.
  • The Goal: We want a drummer who knows what to play and hits it on the beat, all by just watching the conductor.

4. The Conclusion: They Are "Script Readers," Not "World Simulators"

The authors conclude that current AI models are not building a true "physics engine" inside their brains. They aren't simulating how materials vibrate or how sound travels through a room. Instead, they are "script readers" that rely on text descriptions to fill in the gaps.

If you take away the text, the model's ability to understand the physical world (like material hardness or room echo) collapses. The paper argues that for AI to truly understand the world, it needs to learn physics directly from pixels (images), not just from reading captions.

In short: Current video-to-AI audio models are great at making things sound plausible if you give them a hint, but they are terrible at understanding why things sound the way they do when they have to figure it out on their own.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →