← Latest papers
💻 computer science

CapTalk: Text-Guided Stylization and Speech-Driven 3D Head Animation

CapTalk is a novel text-guided framework that enables real-time, synchronized 3D head animation with separate and dynamic control over speaking style and emotional expression by leveraging a large-scale dataset of textual descriptions alongside driving audio.

Original authors: Xuangeng Chu, Yuan Gan, Ziteng Cui, Shuhong Liu, Jian Wang, Bing Zhou, Tatsuya Harada

Published 2026-05-29
📖 4 min read☕ Coffee break read

Original authors: Xuangeng Chu, Yuan Gan, Ziteng Cui, Shuhong Liu, Jian Wang, Bing Zhou, Tatsuya Harada

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have a digital puppet, a 3D head that can talk. Usually, when you feed it audio (like a recording of someone speaking), the puppet just mimics the mouth movements. It's like a ventriloquist's dummy that only moves its lips. If you want the puppet to look happy, angry, or to bob its head in a specific way, you usually have to manually program those actions or feed it a video of a real person doing it first.

CapTalk is a new system that changes the game. Instead of just listening to the audio, it also listens to text instructions you type out. Think of it as a director giving notes to an actor. You can tell the digital head, "Speak this sentence, but do it with a nervous, fast-paced style and a happy expression," or "Say it slowly, with a serious tone and big head nods."

Here is a breakdown of how it works, using simple analogies:

1. The Problem: The "One-Size-Fits-All" Puppet

Previous methods were like a music box. You put in a song (the audio), and it played the exact same tune (the facial movements) every time. If you wanted a different "style," you had to swap out the entire music box mechanism or find a specific recording of a person who already moved that way. This made it hard to create unique characters or change the mood of a conversation on the fly.

2. The Solution: A New "Script" and a New "Library"

The researchers built two main things to solve this:

  • The New Library (The Dataset): They created a massive collection of 200 hours of videos from YouTube. But they didn't just save the video; they used smart AI tools to write a "script" for every clip.

    • One AI listened to the audio to guess the emotion (e.g., "Is this person angry or happy?").
    • Another AI watched the video to describe the physical style (e.g., "Is the mouth opening wide? Is the head nodding a lot?").
    • This created a huge library where every movement is tagged with both an emotion and a style description.
  • The New Director (The Model): They built a model called CapTalk that acts like a translator. It takes three things:

    1. The Audio: What is being said.
    2. The Style Caption: A text description of how to say it (e.g., "subtle head movements," "wide mouth opening").
    3. The Emotion Caption: A text description of the feeling (e.g., "neutral," "excited").

3. How It Works: The "Time-Window" Puzzle

Imagine trying to draw a long, continuous cartoon animation. If you try to draw the whole thing at once, it gets messy. CapTalk breaks the animation into small "time windows" (like short clips of 4 seconds).

  • The Codec (The Shrink Ray): First, the system takes the complex 3D movements of a face and compresses them into a simple code, like turning a high-definition movie into a set of digital instructions (0s and 1s).
  • The Autoregressive Generator (The Storyteller): The model then predicts the next set of instructions based on the previous ones, the audio, and your text notes. It's like a storyteller who knows the plot (the audio) and the character's personality (your text notes), so they can improvise the next few sentences of the story (the facial movement) perfectly.
  • The Magic of Text: If you change the text note from "nervous" to "confident" while keeping the audio the same, the model instantly changes the head movements and mouth shapes to match the new instruction, even in the middle of a sentence.

4. The Results: A Better Performance

The researchers tested CapTalk against other methods and found:

  • Better Lip Sync: The mouth moves perfectly with the words.
  • Better Style Control: If you ask for "big head movements," the head actually moves a lot. If you ask for "subtle," it stays still.
  • Realism: In tests where humans watched the videos, they preferred CapTalk over other methods because the movements felt more natural and matched the instructions better.

In Summary

CapTalk is like giving a digital actor a script that includes not just the dialogue, but also the stage directions. You can type "Speak with a dramatic, wide-eyed style," and the 3D head will instantly perform that specific style, synchronized perfectly with the audio. It moves away from rigid, pre-programmed animations to flexible, text-guided performances.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →