← Latest papers
💻 computer science

Unison: Harmonizing Motion, Speech, and Sound for Human-Centric Audio-Video Generation

Unison is a unified framework that achieves state-of-the-art human-centric audio-video generation by explicitly harmonizing motion, speech, and sound effects through semantic-guided audio decoupling and bidirectional cross-modal forcing strategies to ensure precise temporal alignment and acoustic clarity.

Original authors: Shihao Cheng, Jiaxu Zhang, Quanyue Song, Shansong Liu, Zhizhi Guo, Xiaolei Zhang, Chi Zhang, Xuelong Li, Zhigang Tu

Published 2026-05-12
📖 4 min read☕ Coffee break read

Original authors: Shihao Cheng, Jiaxu Zhang, Quanyue Song, Shansong Liu, Zhizhi Guo, Xiaolei Zhang, Chi Zhang, Xuelong Li, Zhigang Tu

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are directing a movie scene where an actor is singing a song while playing a guitar in the rain. In a perfect world, the singing, the guitar strumming, and the sound of the rain would all happen at the exact right moment, blending together into a beautiful, natural experience.

However, current AI tools trying to create these videos often struggle. They might make the actor's voice so loud it drowns out the guitar and the rain, or they might make the actor's lips move to the guitar notes instead of the words. It's like a band where the singer is screaming over the instruments, and the drummer is playing out of time.

The paper introduces a new AI framework called Unison (which means "acting together") to fix these problems. Think of Unison as a highly skilled audio-visual conductor that ensures every element of the video works in perfect harmony.

Here is how it solves the two main problems, using simple analogies:

1. The "Volume Knob" Problem (Speech vs. Sound Effects)

The Issue: In previous AI models, the "speech" (the actor talking or singing) was like a bully at a party. It would take up all the space, pushing the "sound effects" (like rain, wind, or guitar strums) into the background until they were barely audible. The result was flat, boring audio where the environment felt fake.

The Unison Solution: Unison acts like a smart sound engineer with two separate mixing boards.

  • Separation: Instead of trying to mix the voice and the guitar together in one messy pile, Unison builds them on two separate tracks.
  • The "Smart Gating" (SCG): Imagine a traffic light system. If the scene is mostly about the actor talking, the system automatically turns down the volume of the background noise just enough so the voice is clear. But if the scene is a concert or a storm, the system turns up the background sounds so they don't get drowned out.
  • The "Handshake" (Bi-ACA): The voice track and the sound track constantly "talk" to each other. They check in to make sure the voice isn't accidentally swallowing the guitar sound, and vice versa. This ensures the final audio is a balanced, rich mix where you can hear both the singer and the rain clearly.

2. The "Lip-Sync" Problem (Motion vs. Sound)

The Issue: Often, the video and the audio are out of step. The actor might open their mouth a split second after the sound comes out, or their hand might strum the guitar before the note is heard. It feels like the video and audio are two different people who haven't rehearsed together.

The Unison Solution: Unison uses a training drill called "Cross-Modal Forcing."

  • The Analogy: Imagine two dancers learning a routine. Usually, they try to learn the whole dance at the exact same speed. If one trips, the other trips too, and they get confused.
  • The Unison Twist: Unison lets one dancer (say, the audio) learn the steps slightly faster and cleaner than the other (the video). The "cleaner" dancer then acts as a guide, showing the "noisier" dancer exactly where to step next.
  • The Result: By letting the clearer signal guide the messier one, the AI learns to lock the video movements and the sounds together perfectly. Over time, the actor's lips move exactly when the words are spoken, and the guitar strums hit the exact moment the fingers touch the strings.

The Final Result

When you put these two solutions together, Unison creates videos that feel real.

  • No more drowning out: You can hear the singer and the background music.
  • No more lag: The actions and sounds happen at the exact same time.

The paper shows that Unison doesn't just make videos that look good; it makes them sound good and feel synchronized, creating a much more immersive experience than previous AI models. It achieves this without needing a massive computer, proving that smart organization (like separating the tracks and guiding the learning) is just as important as raw power.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →