← Latest papers
🤖 AI

CETalk: Continuous Valence-Arousal Control for Audio-Driven 3D Talking Head Generation

This paper introduces CETalk, an audio-driven 3D talking head generation framework that leverages continuous Valence-Arousal representations and a multi-scale temporal modeling architecture to achieve fine-grained emotional control and accurate lip synchronization while overcoming the limitations of discrete emotion categories.

Original authors: Peng Jia, Li Dai, Zhen Xiao, Xueliang Liu, Jia Li

Published 2026-08-18
📖 4 min read☕ Coffee break read

Original authors: Peng Jia, Li Dai, Zhen Xiao, Xueliang Liu, Jia Li

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

In the rapidly evolving world of digital media, the ability to create lifelike virtual humans has moved from the realm of science fiction into practical application. These digital avatars are increasingly used as virtual assistants, in virtual reality communication, and for interactive entertainment. A central challenge in this field is making these characters speak with their faces, not just their mouths. While early systems could successfully match lip movements to spoken words, they often struggled to convey the subtle, shifting emotions that make human conversation feel real. Traditional approaches treated emotions as distinct, separate boxes, such as "happy," "sad," or "angry." This method worked for broad strokes but failed to capture the fluid, continuous nature of human feeling, where a smile might slowly fade into a look of concern or where excitement might simmer quietly before bursting forth. Furthermore, the timing of speech and the timing of emotion are fundamentally different; the mouth must move quickly to match the rapid sounds of words, while emotional expressions often evolve more slowly across the entire face.

Researchers at the Hefei University of Technology have developed a new system called CETalk to solve these specific problems. Instead of forcing emotions into rigid categories, this system uses a continuous two-dimensional map to describe feelings, tracking how positive or negative a feeling is and how active or passive it feels. This approach allows the computer to generate facial animations that shift smoothly, mirroring the natural ebb and flow of human conversation. The team built a large new dataset to train their system, reconstructing 3D facial movements from existing video recordings and automatically estimating these continuous emotional values for every single frame. This data provided the necessary foundation for teaching the computer to distinguish between the fast, precise movements required for speech and the slower, sweeping changes that convey mood.

The core of the new system lies in how it handles time. The researchers realized that speech and emotion operate on different speeds, much like how a drummer keeps a steady, fast beat while a singer holds a long, slow note. To manage this, the system splits the processing into two parallel paths. One path focuses on the high-speed details of the mouth and jaw, ensuring that every syllable is synchronized perfectly with the audio. The other path focuses on the slower, broader movements of the face that express emotion, such as a furrowed brow or a widening of the eyes. These two streams of information are then brought back together by a smart integration system that decides, moment by moment, how much weight to give to the speech movements versus the emotional expression. This allows the final animation to be both accurate in its lip-syncing and rich in its emotional depth.

To test their work, the team compared their system against several existing methods using three different sets of video data. The results showed that their approach produced significantly more accurate lip movements and more realistic facial expressions than the previous best methods. In tests on the MEAD dataset, their system reduced lip movement errors to approximately 7.48 millimeters and overall facial movement errors to about 9.91 millimeters, outperforming all other tested technologies on that specific benchmark. Beyond just accuracy, the system demonstrated a superior ability to follow continuous emotional instructions. When asked to generate a face that gradually shifted from a negative state to a positive one, the system followed the path with high precision, matching the intended emotional trajectory far better than its competitors.

The researchers also demonstrated that the system allows for fine-grained control over the intensity of an emotion. By adjusting the input values, they could make a character's expression subtly more intense or more subdued without breaking the synchronization with the voice. This level of control is crucial for creating digital humans that feel truly alive, capable of the nuanced, continuous emotional shifts that define real human interaction. By moving away from discrete categories and embracing the continuous nature of human affect, this work offers a significant step forward in making virtual communication feel less like watching a machine and more like connecting with a person.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →