← Latest papers
💻 computer science

Native Audio-Visual Alignment for Generation

NAVA is a novel framework for joint audio-video generation that employs a native audio-visual alignment strategy within an Align-then-Fuse MMDiT architecture to achieve superior synchronization, video quality, and controllable speech timbre while using only 6.3B parameters.

Original authors: Longbin Ji, Guan Wang, Xuan Wei, Chenye Yang, Xiangrui Liu, Zhenyu Zhang, Shuohuan Wang, Yu Sun, Jingzhou He

Published 2026-05-29
📖 4 min read☕ Coffee break read

Original authors: Longbin Ji, Guan Wang, Xuan Wei, Chenye Yang, Xiangrui Liu, Zhenyu Zhang, Shuohuan Wang, Yu Sun, Jingzhou He

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you want to create a movie scene where a character is riding a bicycle while talking. You need two things to happen perfectly together: the video of the bike moving and the sound of the wind and the person's voice.

For a long time, computers tried to make these two things separately and then glued them together at the end. This is like having one artist draw the bike and a completely different artist record the voice, then trying to make them match up later. Often, the voice would be slightly out of sync with the lip movements, or the sound wouldn't feel like it belonged in that specific scene.

Other newer methods tried to put the artist, the voice actor, and the director all in the same tiny room to work on everything at once. While this helps them talk to each other, it gets messy. The "director" (the text instructions) ends up getting confused with the "sound engineer" (the timing of the noise), making it hard to control specific details like who is speaking or what their voice sounds like.

Enter NAVA: The "Rehearsal Room" Approach

The paper introduces NAVA (Native Audio-Visual Alignment), which changes the workflow. Instead of working separately or getting everything mixed up in one big pot, NAVA uses a clever two-step process called "Align-then-Fuse."

Think of it like a theater rehearsal:

  1. The Rehearsal (Alignment): First, the actors (audio) and the stage crew (video) meet in a dedicated "rehearsal room." Here, they practice together without the director giving them new lines. They figure out exactly how the sound of a footstep matches the visual of a foot hitting the ground, or how a guitar strum matches a hand movement. They build a natural, native connection between sound and sight.
  2. The Performance (Fusion): Once they know how to move and speak together, the director (the text prompt) steps in. The director doesn't need to teach them how to sync up anymore; they just tell them what to do. "Now, the character should say this line with a grumpy voice." Because the actors and crew already know how to move in sync, they can instantly adapt to the new instructions without losing their timing.

The Secret Sauce: "Timbre-in-Context"

NAVA also has a special trick for handling multiple speakers. Imagine a script where a mother and her child are talking. You don't want the mother's voice to sound like the child's.

Older methods might have a separate "voice switch" button that changes the whole scene's voice. NAVA is smarter. It treats the voice style (timbre) as part of the script itself. It attaches a specific "voice tag" to the mother's lines and a different "voice tag" to the child's lines. When the computer generates the scene, it knows exactly which voice belongs to which sentence, just like a director reading a script with notes on who is speaking.

What the Paper Found

The researchers tested NAVA against other top models using a benchmark called Verse-Bench (a test of how well audio and video match) and Seed-TTS (a test of voice control).

  • Better Sync: NAVA was the best at making sure the sound happened at the exact moment the visual event occurred (like a door slamming or lips moving).
  • High Quality: The video looked great, and the audio sounded clear and realistic.
  • Efficiency: Despite being smaller than many competitors (using only 6.3 billion "brain cells" or parameters), it outperformed much larger models.
  • Control: It could switch between different speakers in the same scene seamlessly, keeping the voices distinct and accurate.

In Summary

NAVA solves the problem of making sound and video by giving them a dedicated space to learn how to move together before they try to follow complex instructions. It's like teaching a dance duo to master their steps before the choreographer tells them which song to dance to. The result is a movie scene where the sound and picture feel like they were born together, not just glued together.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →