← Latest papers
💻 computer science

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections

To address reproducibility and copyright challenges in video-to-music generation, this paper introduces the reproducible Open Screen Soundtrack Library version 2 (OSSL-v2) dataset of public-domain films and proposes a dialogue-aware model that enhances music generation by modulating video cross-attention with on-screen speech.

Original authors: Haven Kim, Zachary Novack, Julian McAuley, Hao-Wen Dong

Published 2026-08-13
📖 4 min read☕ Coffee break read

Original authors: Haven Kim, Zachary Novack, Julian McAuley, Hao-Wen Dong

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are watching a movie. You see a hero running through a storm, and suddenly, thunder rumbles while a dramatic orchestra swells. That music isn't just background noise; it's the emotional glue holding the scene together, telling you how to feel. For a long time, scientists have been trying to teach computers to be the "composer" for videos, automatically creating music that matches the mood of what's happening on screen. This field is called video-to-music generation. But there's a huge problem: to teach a computer this skill, you need thousands of examples of movies and their perfect soundtracks. Most researchers have been using lists of links to videos on the internet to find these examples. The trouble is, internet links are like sandcastles; the tide comes in, and the videos get deleted or locked away. This means the "textbooks" the computers learn from keep disappearing, making it impossible to check if one computer is actually smarter than another. To fix this, we need a library of movies that can't vanish, and a way to teach the computer not just what the characters look like, but also what they are saying.

Enter the researchers from UC San Diego and the University of Michigan, who have built a brand-new, unshakeable library called OSSL-v2. Think of this as a massive, self-contained movie theater that no one can shut down. Instead of pointing to links on the internet that might rot away, they downloaded 1,886 films that belong to the public (meaning anyone can use them) and chopped them into 34,343 little clips. This adds up to about 246.4 hours of video and music, all safely stored in one place. Because it's self-hosted, it won't disappear, and it gives every researcher a fair, identical set of data to train their AI models on.

But the team didn't just build a library; they realized the computers were missing a crucial clue. In movies, the music often changes the moment a character starts speaking. If someone is shouting, the music might get quieter or stop to let the voice be heard. The researchers noticed that in their new library, there is a slight but real connection between how loud the dialogue is and how loud the music is. To use this, they invented a "dialogue-aware" system. Imagine the computer is a conductor. Before, the conductor only looked at the actors' faces and movements to decide when to speed up or slow down the orchestra. Now, the conductor has a new pair of glasses that lets them hear the actors' voices in real-time. The system takes the sound of the dialogue and uses it to gently nudge the music generation frame-by-frame, ensuring the music breathes and pauses exactly when the characters do.

When they tested this new approach on their massive library, the results were promising. They compared their "dialogue-aware" models against the best existing systems. The new method didn't necessarily make the music sound "better" in a general sense (like making it more complex or diverse), but it did make the music fit the video much more tightly. Specifically, the generated soundtracks became more faithful to the original movie scenes, especially when the computer had to handle tricky moments where dialogue and music overlap. The researchers found that this improvement held true even when they tested the models on commercial movies they had never seen before, suggesting that listening to the dialogue helps the AI understand the rhythm of a scene better than just watching it.

The paper also carefully ruled out a few potential tricks. They worried that maybe the computer was just "relying on" leftover music leaking into the dialogue track (since separating audio isn't perfect). To check this, they ran a special filter to find clips where the dialogue was almost completely free of any music noise. Even with these "clean" clips, the dialogue-aware system still performed better. This suggests the improvement really comes from understanding the speech, not from accidentally hearing the music. While the system isn't a perfect replacement for human composers yet, the study suggests that adding the voice of the characters is a powerful way to help computers write music that feels truly connected to the story.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →