← Latest papers
💻 computer science

EchoFoley: Event-Centric Hierarchical Control for Video Grounded Creative Sound Generation

To overcome key limitations in existing video-text-to-audio models, the paper introduces EchoFoley, a new event-centric task with hierarchical control, supported by the EchoFoley-6k benchmark and the EchoVidia framework, which significantly improves both controllability and perceptual quality in video-grounded sound generation.

Original authors: Bingxuan Li, Yiming Cui, Yicheng He, Yiwei Wang, Shu Zhang, Longyin Wen, Yulei Niu

Published 2026-06-24
📖 4 min read☕ Coffee break read

Original authors: Bingxuan Li, Yiming Cui, Yicheng He, Yiwei Wang, Shu Zhang, Longyin Wen, Yulei Niu

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have a silent movie playing on a screen. You can see a cat walking, a door slamming, and a car driving by. Now, imagine you want to add sound effects, but not just any sound. You want the cat to meow gently at first, then suddenly roar like a lion when a wizard casts a spell, and you want that specific roar to happen exactly at the 7-second mark, while making all the sounds before it louder than the sounds after it.

Current AI tools are like a clumsy sound engineer who hears "cat" and just dumps a generic "meow" sound file over the whole video. They struggle to listen to your specific, detailed instructions.

EchoFoley is a new project designed to fix this. Here is how it works, broken down into simple concepts:

1. The Problem: The "Visual Dominance" Trap

Right now, if you tell an AI, "Make the second meow sound like a lion," the AI often gets confused. It sees the cat (the visual) and thinks, "Okay, I'll just make a cat sound." It ignores your specific text instructions because it relies too heavily on what it sees rather than what you say. It's like a chef who only cooks what they see on the plate, ignoring your request to "add more salt."

2. The Solution: A "Sound Script" (Symbolic Representation)

The researchers created a new way to talk to the AI. Instead of just giving a vague command, they teach the AI to write a "Sound Script."

Think of this script like a musical conductor's score. It doesn't just say "play music"; it breaks the sound down into tiny, specific notes:

  • When: Exactly what second does the sound happen?
  • What: Is it a cat meow or a lion roar?
  • How: Is it loud? Is it high-pitched? Is it coming from the left or right?

By forcing the AI to write this script first, it can handle complex requests like, "Change the second meow to a lion roar, but keep the first one normal."

3. The New Playground: EchoFoley-6k

To teach the AI this new skill, the team built a massive training library called EchoFoley-6k.

  • Imagine a library with 6,000 silent videos.
  • For each video, they didn't just write one sentence; they wrote 6,000 detailed instructions and 42,000 tiny sound notes.
  • They hired experts to label exactly when a sound starts and stops, and what properties it should have. This is the "textbook" the AI learns from.

4. The New Brain: EchoVidia (The "Slow-Fast" Thinker)

The team built a new AI system called EchoVidia to use this library. It uses a clever trick called "Slow-Fast Thinking," inspired by how humans think:

  • Fast Thinking (System 1): The AI glances at the video quickly (1 frame per second) to get the general vibe. "Oh, it's a cat video."
  • Slow Thinking (System 2): The AI then slows the video down to a crawl (watching it in slow motion) to look closely. "Wait, I see the cat's mouth open at 00:04. That's when the meow happens. And at 00:07, the wizard wave happens."

By combining a quick overview with a slow, detailed inspection, the AI can pinpoint exactly when to put a sound and what that sound should be, rather than just guessing based on the general scene.

5. The Results: A Masterful Sound Engineer

When they tested EchoVidia against other top AI models:

  • Control: It was 40% better at following specific instructions. If you asked for a sound at a specific time, it actually did it.
  • Quality: It sounded 12% more natural and realistic to human listeners.
  • Balance: Unlike other models that ignored your text instructions to focus on the video, EchoVidia successfully listened to both the video and your specific commands.

In Summary

The paper introduces a new way to make AI generate sound for videos. Instead of letting the AI guess based on the picture, they gave it a detailed script and a slow-motion thinking process to ensure every sound happens at the right time, with the right tone, exactly as the user requested. It turns a clumsy, guess-and-check process into a precise, creative tool for storytelling.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →