← Latest papers
💻 computer science

UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos

UniSwap is a novel framework that enables the first streaming, joint audio-visual identity replacement in talking videos by utilizing a single diffusion transformer with a swap-and-reconstruct training pipeline and efficient sampling techniques to achieve synchronized, consistent, and stable long-form generation.

Original authors: Yuxuan Zhang, Haozhong Xiong, Jiayi Song, Jinpeng Yu, Yang Shi, Jiaming Liu, Ruihua Huang, Liwei Wang

Published 2026-08-13
📖 4 min read☕ Coffee break read

Original authors: Yuxuan Zhang, Haozhong Xiong, Jiayi Song, Jinpeng Yu, Yang Shi, Jiaming Liu, Ruihua Huang, Liwei Wang

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are watching a movie, but the actor on screen suddenly starts speaking with a different voice and looking like a completely different person, yet they are still saying the exact same words and moving their lips in perfect time. This is the dream of "identity swapping" in talking videos. For a long time, computers have been great at doing just one part of this trick: either changing a person's face while keeping their original voice, or changing a voice while keeping the original face. But trying to do both at the same time has been like trying to juggle two different balls while riding a unicycle; the results were often out of sync, with lips moving to the wrong words or voices sounding robotic. This paper dives into the world of generative AI, specifically looking at how we can teach computers to swap both a person's look and their voice simultaneously, in real-time, without the video stuttering or the audio lagging behind.

The researchers behind this project, UniSwap, have built a new system that acts like a master puppeteer, pulling the strings for both the visual and audio sides of a video at the exact same moment. Instead of using two separate tools—one for the face and one for the voice—they created a single "brain" that understands how a person's face moves and how their voice sounds as one connected experience. Think of it like a dual-language translator who doesn't just translate words but also captures the speaker's accent and facial expressions instantly.

The paper introduces a clever way to teach this system. Since it's nearly impossible to find real-life videos of the same person speaking the same lines while looking and sounding like two different people, the team invented a "swap-and-reconstruct" game. They take a real video, digitally strip the person's face and voice away to create a "ghost" version, and then ask the AI to rebuild the original video using a new face and voice as a guide. It's like giving a chef a blank canvas and a recipe, then asking them to recreate a specific dish perfectly. By training on millions of these "rebuilds," the AI learns to swap identities while keeping the original movements and background intact.

What makes this truly special is that it works like a live stream rather than a slow, pre-recorded movie. Most video generators have to watch the whole clip before they can start drawing the first frame, which is too slow for interactive use. UniSwap, however, generates the video in small blocks, like a conveyor belt, allowing it to produce about 13.6 frames per second on a powerful computer chip (an NVIDIA H100). While this isn't quite fast enough for instant real-time interaction yet, it is a massive leap forward, allowing the system to generate long videos—up to an hour in length—without the character's face or voice drifting apart or becoming distorted over time. The authors found that by using a special "memory trick" called Feature-RoPE Decomposition, the AI can remember the original identity even after generating thousands of frames, ensuring the character stays consistent from the first second to the last.

In short, UniSwap suggests that we can finally have a unified system that swaps both appearance and voice in talking videos with high synchronization and stability. While the current version isn't quite fast enough for live, real-time conversation (it's slightly slower than the speed of a standard video playback), it proves that a single model can handle both tasks together better than trying to stitch two separate systems together. The results show that the lip movements stay perfectly synced with the new voice, and the character's identity remains steady even in long clips, offering a promising step toward more natural and interactive digital avatars.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →