← Latest papers
💻 computer science

Omni-Customizer: End-to-End MultiModal Customization for Joint Audio-Video Generation

Omni-Customizer is an end-to-end framework that enables precise joint audio-video generation with consistent visual identities and vocal timbres across multiple subjects by introducing novel modules like Omni-Context Fusion, Masked TTS Cross-Attention, and Semantic-Anchored Multimodal RoPE to solve challenges like speech leakage and achieve state-of-the-art multimodal customization.

Original authors: Yuheng Chen, Qingdong He, Teng Hu, Yuji Wang, Yabiao Wang, Lizhuang Ma, Jiangning Zhang

Published 2026-05-19
📖 4 min read☕ Coffee break read

Original authors: Yuheng Chen, Qingdong He, Teng Hu, Yuji Wang, Yabiao Wang, Lizhuang Ma, Jiangning Zhang

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are a director trying to film a movie scene with multiple actors. You have a script, but you also need to make sure that Actor A always looks like the specific person you hired and sounds exactly like their unique voice, while Actor B does the same for themselves. If the camera gets confused, you might end up with Actor A's face moving while Actor B's voice is speaking, or worse, the actors might start speaking lines that aren't in the script just because you described them in the scene notes.

This is the exact problem Omni-Customizer solves. It is a new AI system designed to create videos where multiple people can talk and interact, while perfectly keeping their unique looks and voices intact.

Here is how it works, broken down with simple analogies:

1. The Problem: The "Confused Director"

Previous AI tools were good at making videos or good at making audio, but when they tried to do both at once with multiple people, they got confused.

  • The Mix-up: The AI might mix up who is speaking. It might make the man in the red shirt speak the woman's lines.
  • The "Ghost" Voice: Sometimes, the AI would accidentally turn the description of the scene into speech. If you wrote, "The man is tall," the AI might make the character actually say, "I am tall," even though that wasn't in the dialogue.
  • The Drift: As the video goes on, the characters might slowly change their faces or voices, losing their identity.

2. The Solution: The "Omni-Customizer" Toolkit

The researchers built a system with three main "tools" to fix these issues:

A. The "Super-Connector" (Omni-Context Fusion)

Think of the AI's brain as having different departments: one for text, one for pictures, and one for sound. Usually, these departments talk to each other very slowly and indirectly.

  • The Fix: Omni-Customizer builds a "Super-Connector" that forces all these departments to sit at the same table and talk instantly. It takes the description of the person, their photo, and their voice sample, and fuses them together into one tight package before the video generation even starts. This ensures the AI knows, "This face, this voice, and this name all belong to the same person."

B. The "Name Tag" System (Semantic-Anchored RoPE)

Imagine you have a pile of puzzle pieces: some are faces, some are voices, and some are words. If you just throw them in a box, they get mixed up.

  • The Fix: The system uses a special "Name Tag" (called SA-MRoPE). It physically attaches the puzzle piece of "Face A" and "Voice A" directly to the text word "Subject 1" in the script. It's like putting a magnetic tag on the puzzle pieces so they cannot float away and attach to the wrong person. This stops the AI from getting confused about who is who, even in crowded scenes with three or four people.

C. The "Silence Button" (Masked TTS Cross-Attention)

Remember the problem where the AI accidentally spoke the scene descriptions?

  • The Fix: The researchers installed a "Silence Button" (MTP-CA). They tell the AI: "Only speak the words inside the <S> (Start) and <E> (End) tags." Everything else—like descriptions of the weather or the characters' clothes—is strictly silenced. The AI is forced to ignore the descriptive text when generating speech, so it never accidentally reads the stage directions out loud.

3. The Training: "The Gym Routine"

To teach this system, the researchers didn't just throw all the data at it at once. They used a smart training schedule:

  • The "Audio-Only" Drills: Sometimes, the AI practices only with audio (listening to voices and learning languages) without worrying about the video. This helps it learn to speak many different languages without getting confused.
  • The "Video-Audio" Drills: Other times, it practices making the video and audio match perfectly (lip-syncing).
  • The "Easy to Hard" Ladder: They started by teaching the AI to handle just one person talking. Once it mastered that, they moved to two people, and finally to complex scenes with multiple people talking over each other. This step-by-step approach prevented the AI from getting overwhelmed.

4. The Result

When tested, Omni-Customizer proved it could:

  • Keep faces and voices consistent, even in long conversations.
  • Prevent the "Ghost Voice" problem (where descriptions are spoken).
  • Handle complex scenes with multiple people better than any previous open-source model.

In short: Omni-Customizer is like a highly organized film crew that never loses track of which actor is which, ensures they only speak their actual lines, and keeps their voices and faces consistent from the first second to the last.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →