← Latest papers
💻 computer science

UniLS: End-to-End Audio-Driven Avatars for Unified Listening and Speaking

UniLS is the first end-to-end framework that generates unified speaking and listening avatar expressions driven solely by dual-track audio, overcoming the stiffness of previous listener models through a novel two-stage training paradigm that learns an internal motion prior before modulating it with speech cues.

Original authors: Xuangeng Chu, Ruicong Liu, Yifei Huang, Yun Liu, Yichen Peng, Bo Zheng

Published 2026-03-31
📖 4 min read☕ Coffee break read

Original authors: Xuangeng Chu, Ruicong Liu, Yifei Huang, Yun Liu, Yichen Peng, Bo Zheng

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are at a lively dinner party. You have two goals: you want to tell a great story (speaking), but you also want to show your friends you are listening when they talk (listening).

For a long time, computer scientists could build digital avatars that were great at telling stories. But when it came to listening, these avatars were terrible. They would just sit there like statues, staring blankly, blinking once every ten minutes, and looking completely bored. They were "stiff."

This paper introduces UniLS, a new system that finally teaches digital avatars how to be good conversationalists—both speakers and listeners.

Here is how they did it, explained simply:

The Problem: The "Stiff Listener"

The researchers realized that teaching a computer to speak is easy because the sound of the voice directly controls the mouth. If you say "Hello," your lips move. It's a direct link.

But listening is different. When someone else is talking, your face doesn't just copy their words. You nod, you blink, you raise an eyebrow, or you smile. These reactions come from inside you, not just from the sound of their voice.

When previous AI tried to learn this by just listening to audio, it got confused. It thought, "If there is no sound coming from me, I should do nothing." So, the avatar froze.

The Solution: A Two-Step Dance Lesson

To fix this, the team created a two-step training process, like teaching a dancer in two different phases.

Step 1: The "Silent Rehearsal" (Learning the Inner Rhythm)

First, they taught the AI to move without any audio at all.

  • The Analogy: Imagine a dancer practicing alone in a room with no music. They aren't trying to match a beat; they are just practicing how to breathe, how to shift their weight, how to blink naturally, and how to nod when they are thinking.
  • What the AI learned: The AI learned a "motion library" of natural, spontaneous human movements. It learned that humans blink often, tilt their heads, and fidget slightly even when they aren't speaking. This is called the Internal Motion Prior.

Step 2: The "Duet" (Adding the Music)

Once the AI knew how to move naturally on its own, they introduced the audio.

  • The Analogy: Now, the dancer is on stage with a partner. The partner is speaking (the music). The dancer uses their "silent rehearsal" skills (the natural blinking and nodding) but adjusts them based on the partner's voice. If the partner laughs, the dancer smiles. If the partner pauses, the dancer leans in.
  • What the AI learned: The AI learned to take those natural, internal movements and gently nudge them to match the conversation. It didn't stop moving; it just made the movement react to the other person.

Why This Matters

Before this, digital avatars were like actors who could only deliver a monologue. If you asked them to listen, they turned into mannequins.

UniLS is the first system that is End-to-End. This means you just give it two audio tracks (Person A talking, Person B talking), and it instantly generates the faces for both people.

  • Person A talks with perfect lip-sync.
  • Person B listens with natural nods, blinks, and expressions.

The Results

The researchers tested this against other methods, and the results were huge:

  • Speaking: The avatars spoke just as well as the best existing systems.
  • Listening: The avatars improved by 44% in naturalness. They stopped looking like statues and started looking like real people having a conversation.

The Bottom Line

UniLS is like giving a digital human a "social brain." It understands that conversation isn't just about talking; it's about the silent, subtle dance of listening, too. Now, when you talk to these avatars, they don't just hear you; they see you, and they react like a real human would.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →