← Latest papers
🤖 AI

OmniNFT: Modality-wise Omni Diffusion Reinforcement for Joint Audio-Video Generation

OmniNFT is a novel modality-aware online diffusion reinforcement learning framework that enhances joint audio-video generation by addressing multi-objective inconsistency, gradient imbalance, and uniform credit assignment through modality-wise advantage routing, layer-wise gradient surgery, and region-wise loss reweighting.

Original authors: Guohui Zhang, XiaoXiao Ma, Jie Huang, Hang Xu, Hu Yu, Siming Fu, Yuming Li, Zeyue Xue, Lin Song, Haoyang Huang, Nan Duan, Feng Zhao

Published 2026-05-13
📖 4 min read☕ Coffee break read

Original authors: Guohui Zhang, XiaoXiao Ma, Jie Huang, Hang Xu, Hu Yu, Siming Fu, Yuming Li, Zeyue Xue, Lin Song, Haoyang Huang, Nan Duan, Feng Zhao

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to teach a robot to create a movie where the sound and the picture are perfectly in sync. You want the robot to learn from its mistakes, just like a human does. This is what the paper calls "Reinforcement Learning."

However, the authors found that when they tried to teach the robot to make both audio and video at the same time using standard methods, the robot got confused. It was like trying to coach a basketball player and a singer simultaneously, but the coach was shouting contradictory instructions.

Here is the story of OmniNFT, the new method the authors built to fix this, explained through simple analogies.

The Three Big Problems

The authors discovered three specific reasons why the robot was failing:

  1. The "Confused Coach" Problem (Advantage Inconsistency):
    Imagine a coach grading a student's performance. Sometimes, the student's video is amazing, but the audio is terrible. In standard training, the coach might give a single "average" grade for the whole performance.

    • The Issue: If the video is great but the audio is bad, an average grade might tell the robot, "You did okay," which doesn't help it fix the audio. Or worse, the robot might think it needs to change the video to fix the audio, making the video worse.
    • The Fix: OmniNFT acts like a coach with two separate scorecards. One scorecard judges the video, and the other judges the audio. If the video is good, the video part of the robot gets a "high five." If the audio is bad, the audio part gets a "time out." They don't mix the scores up.
  2. The "Noisy Classroom" Problem (Gradient Imbalance):
    Think of the robot's brain as a multi-story building.

    • The bottom floors (shallow layers) are where the audio is generated (making the voice sound right).
    • The top floors (deep layers) are where the video and audio talk to each other to stay in sync.
    • The Issue: When the robot tries to learn from the video, the "noise" from the video lessons was accidentally leaking down into the bottom floors, messing up the audio generation. It was like the video teacher shouting instructions in the audio classroom, confusing the audio students.
    • The Fix: The authors installed a soundproof wall (Gradient Surgery). They let the video and audio talk to each other on the top floors where they need to sync up, but they blocked the video noise from reaching the bottom audio floors. This lets the audio learn to be clear without being distracted by the video.
  3. The "Spotlight" Problem (Uniform Credit Assignment):
    Imagine a scene where a character is speaking. The most important part of the video is their mouth moving, and the most important part of the audio is their voice.

    • The Issue: Standard training treats every pixel and every sound wave equally. It's like shining a giant, flat floodlight over the whole stage. The robot spends just as much time learning about the background wall as it does learning about the actor's lips.
    • The Fix: OmniNFT uses a smart spotlight. It looks at the video and sees exactly where the sound is coming from (like a person's mouth). It then shines a bright, intense light on just those specific areas. This tells the robot, "Pay extra attention here! This is where the magic happens."

The Result: A Better Movie Maker

By using these three tricks, the authors tested their new system (called OmniNFT) on a powerful existing model called LTX-2.

The results were impressive:

  • Better Quality: The videos looked sharper, and the sounds were clearer.
  • Better Sync: The lips moved exactly when the words were spoken.
  • Better Harmony: The video and audio felt like they belonged together, rather than two separate things glued together.

In a Nutshell

The paper says that to make great AI movies, you can't just use a "one-size-fits-all" teaching method. You have to:

  1. Grade the sound and picture separately.
  2. Stop the video lessons from confusing the audio lessons.
  3. Focus extra attention on the parts of the movie where sound and picture actually meet (like a talking face).

OmniNFT does exactly this, resulting in AI-generated videos that look and sound much more realistic and synchronized than before.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →