← Latest papers
💻 computer science

Vorch-Streamer: Extending Human Audio-Visual Generation to Real-Time Long-Form Streaming

Vorch-Streamer is a post-training framework that enables real-time, long-form text-to-audio-video streaming by combining a causal generator trained with mixed forcing strategies, DMD distillation for long-horizon stability, and an external language model for speech planning to achieve high-speed, synchronized, and consistent avatar generation.

Original authors: Menglin Han, Yang Ding, Yulei Lu, Haoran Yu, Xin Ma, Junyi Chen, Zhangkai Ni, Lin Ma, Yaohui Wang

Published 2026-08-07
📖 3 min read☕ Coffee break read

Original authors: Menglin Han, Yang Ding, Yulei Lu, Haoran Yu, Xin Ma, Junyi Chen, Zhangkai Ni, Lin Ma, Yaohui Wang

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to teach a robot to tell a story. You want it to speak, move its lips, and change its facial expressions all at the same time, just like a real person. For a long time, the best robots could only do this for very short clips, like a 5-second greeting. They worked like a director watching the entire script before filming a scene, looking ahead to see what happens next. But in the real world, we don't have time to wait for the whole story to be written before we start talking. We need the robot to speak now, based only on what it has said so far, while keeping its face and voice perfectly in sync. This is the challenge of "real-time streaming." The problem is that if a robot makes a tiny mistake in the first second, that mistake can pile up over time, making the robot's face look weird or its voice drift off-topic by the end of the conversation. Scientists have been trying to figure out how to make these digital avatars talk for minutes at a time without losing their minds or their rhythm.

Enter Vorch-Streamer, a new system that acts like a super-organized, real-time storyteller. Instead of trying to memorize the whole script before speaking, Vorch-Streamer uses a clever two-part strategy to keep the show going for over two minutes without stopping. Think of it like a band playing a long jam session. Usually, if a musician plays a wrong note, the rest of the song might get messy. But Vorch-Streamer has a special "rehearsal" mode. First, it practices playing short segments while being gently corrected by a teacher (a pre-trained model that knows how to do it perfectly). Then, it practices playing the whole song from start to finish, but this time, it listens to its own previous mistakes and learns how to fix them on the fly, all while a "conductor" (an AI planner) tells it exactly which words to say next.

The result is a digital avatar that can generate video and audio simultaneously at 27.12 frames per second (FPS). To put that in perspective, standard video plays at 24 FPS, meaning this robot is actually faster than real-time! It can keep its face looking like the same person, its lips moving in perfect time with its words, and its voice clear and accurate for nearly two minutes straight. The researchers found that without a specific "planning" step to decide what to say next, the robot would get confused and start speaking gibberish or repeating itself. By adding this planner, the system successfully solves the puzzle of how to keep a long conversation flowing smoothly without needing to see the future. It proves that you can take a powerful, slow robot that knows everything about a story and turn it into a fast, live performer that knows just enough to keep the show going, one second at a time.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →