Vorch-IR: Long-Form Unified Multimodal Identity Replacement Video Generation
Vorch-IR is a unified multimodal framework built on LTX2 that enables flexible single- and dual-person identity replacement with optional background editing in long-form videos by jointly conditioning on driving videos, reference images, and textual instructions, while overcoming data scarcity through an automatic synthesis pipeline and extending to minute-long generations via a temporal overlapping inference strategy.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are holding a remote control for reality, but instead of changing the channel, you want to swap the actors in a movie while keeping the script, the camera angles, and the dance moves exactly the same. This is the dream of "video identity replacement," a hot topic in the world of artificial intelligence. For a long time, AI could generate new videos from scratch, but editing an existing one was like trying to change a character's face in a live-action film without messing up the lighting or the movement. Early attempts required the AI to be handed a detailed map of where the person was standing, a skeleton drawing of their pose, or a mask to cover their face. It was like asking a chef to cook a meal only if you gave them a pre-cut list of ingredients and a diagram of the kitchen. These methods were rigid and hard to use. The big question researchers are asking now is: Can we teach an AI to understand a simple sentence like "Swap the dancer on the left with this photo" and just get it, without needing a complex map or a skeleton?
Enter Vorch-IR, a new AI system that acts like a super-smart, magical film editor. Think of it as a director who doesn't need a script supervisor or a stunt coordinator. You give Vorch-IR three things: the original video (the "driving" video), a few photos of the new people you want to insert (the "references"), and a simple text note telling the AI who goes where. The magic is that the photos don't need to match the video's pose or position. If the video shows a person dancing with arms up, and your photo shows them sitting down, Vorch-IR figures out how to make the sitting person dance anyway. It handles single people, two people dancing together, and even swapping the background scenery, all with one single model.
The researchers behind Vorch-IR built this system on top of a powerful video generator called LTX2. They taught it to look at the video, the photos, and the text instructions all at once. Instead of using rigid maps, the AI uses a "vision-language" brain (a type of AI that understands both pictures and words) to figure out the connection. It's like the AI reads your note, looks at the photo, and says, "Ah, you want this person to take the place of the one on the left," and then it uses its internal "self-attention" to blend the new face onto the moving body perfectly.
One of the biggest hurdles in this field is data. To teach an AI to swap faces, you need thousands of examples of "before" and "after" videos. Since these don't exist in the real world, the team built a robot pipeline to make them. They took real videos, used other AI tools to swap the faces in the first frame, and then used a motion-guided robot to make the rest of the video follow the new faces. They then filtered out the bad attempts (like when the AI accidentally gave the dancer three arms) to create a perfect training set. This allowed them to train one model to do everything: swapping one person, swapping two people, and even changing the background, all without needing separate tools for each job.
But what about making a whole movie? Most AI video generators get confused or blurry after a few seconds. Vorch-IR uses a clever trick called "temporal overlapping inference." Imagine trying to paint a long mural. Instead of painting one small square, drying it, and then painting the next (which might make the colors look different), Vorch-IR paints overlapping sections at the same time and blends them together seamlessly. This allows it to generate videos that are a full minute long without the characters' faces drifting or the motion getting weird, all without needing to generate the video one clip at a time.
The team tested Vorch-IR against other top AI models. In head-to-head comparisons, Vorch-IR suggested that it was better at keeping the new person's face looking exactly like the photo (identity preservation) and keeping the background consistent, while still moving smoothly. In human tests, people preferred Vorch-IR's results over other methods, especially when it came to making the new character look real and physically plausible. The paper suggests that while some older models might be slightly better at specific things like perfect pose matching in very specific scenarios, Vorch-IR offers a much more flexible and unified way to edit videos, proving that you don't need complex maps to get great results. It's a step toward a future where editing a video is as easy as typing a sentence.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.