← Latest papers
💻 computer science

OmniMate: Open-Ended Real-Time Streaming Audio-Visual Generation for Interactive Avatars

OmniMate is a unified framework that enables high-quality, low-latency, open-ended real-time streaming audio-visual avatar generation by introducing a Generation Progress Controller for adaptive response progression and a Multi-Reference Conditioning Module to ensure long-term cross-modal identity consistency.

Original authors: Quanyue Song, Yishan He, Yanbo Ding, Zhixiang He, Yongxiang Li, Caigui Jiang, Zhi Zhi Guo

Published 2026-07-29
📖 4 min read☕ Coffee break read

Original authors: Quanyue Song, Yishan He, Yanbo Ding, Zhixiang He, Yongxiang Li, Caigui Jiang, Zhi Zhi Guo

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are talking to a friend who lives inside a computer screen. You want them to not just speak, but to move their hands, blink, and react to your voice in real-time, just like a real person. This is the dream of "interactive avatars." For a long time, making these digital characters was like filming a movie: you had to plan every line and every gesture before you could hit "play." But recently, scientists have built "diffusion models," which are like magical paintbrushes that can instantly create high-quality videos and sounds from text. This technology has gotten so good that it can now make a character speak and move in sync with audio. However, there's a catch: these models usually work best when they know exactly how long the conversation will be. If you try to have a free-flowing chat where the computer has to guess when to stop talking and start listening again, the character often gets confused, keeps talking too long, or starts to look like a different person after a few minutes. This paper tackles that exact problem: how to make a digital friend who can chat with you forever, without losing their face or their voice, and without lagging behind your words.

The researchers behind this project, a team from Xi'an Jiaotong University and China Telecom, have built a new system called OmniMate. Think of OmniMate as a super-smart digital puppeteer that can handle a never-ending conversation in real-time. Unlike previous systems that might stumble when the conversation gets long or when the user interrupts, OmniMate is designed to flow naturally, switching instantly between "talking" and "listening" modes while keeping the character looking and sounding exactly the same.

The secret to OmniMate's success lies in two clever tricks it uses to stay on its toes. First, it uses something called a Generation Progress Controller (GPC). Imagine you are reading a storybook, but you don't know how many pages are left. Usually, you might keep reading past the end or stop too early. The GPC is like a built-in page counter that tells the avatar exactly how much of its response is left to generate. It gives the system a "progress bar" for every tiny chunk of video and audio it creates. This allows the avatar to know precisely when to finish a sentence and smoothly switch back to a "listening" pose, waiting for your next input. Without this, the avatar might keep talking after it's done, or freeze awkwardly.

The second trick is the Multi-Reference Conditioning Module (MRCM). Have you ever noticed that if you look at a photo of a person from a weird angle, they might look a little different? In long videos, digital avatars often suffer from "identity drift," where their face slowly morphs into someone else's, or their voice changes pitch. OmniMate solves this by constantly reminding itself who the character is. Instead of just looking at the very first frame of the video, it keeps a "memory bank" of several reference photos and a sample of the character's voice. As the video plays, it constantly checks these references, like a painter glancing back at a portrait to make sure the nose and eyes stay correct, ensuring the character looks and sounds the same whether they are talking for 10 seconds or 10 minutes.

The team tested OmniMate on a benchmark called VerseBench, which simulates real conversations. The results were impressive: the system can generate video and audio at a speed of 27.64 frames per second (FPS), with a "time-to-first-frame" (how long it takes to start showing you the video) of just 3.49 seconds. This is fast enough to feel like a real, live conversation. In tests, OmniMate beat other top models in keeping the character's face and voice consistent over long periods. While other systems might start to look blurry or change their voice after a minute, OmniMate stayed stable even in tests lasting up to 240 seconds.

However, the authors are careful to note that the system isn't perfect yet. If the system guesses the wrong amount of time for a response, the avatar might cut off a word or repeat itself. Also, if the character needs to do something that drastically changes their appearance, like a huge costume change, the system might struggle to keep the identity consistent. But overall, OmniMate suggests a major step forward: a way to have digital friends who can chat with us in real-time, remember who they are, and react to us without the awkward glitches of the past. It turns the static, pre-planned videos of today into dynamic, living conversations for tomorrow.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →