← Latest papers
💻 computer science

Realistic Lip Motion Generation Based on 3D Dynamic Viseme and Coarticulation Modeling for Human-Robot Interaction

This paper presents a lightweight and efficient framework for generating realistic lip motions in humanoid robots by constructing a 3D dynamic viseme library based on Chinese pronunciation theory and employing a coarticulation mechanism with initial-final decoupling to ensure natural synchronization.

Original authors: Sheng Li, Jingcheng Huang, Min Li

Published 2026-04-03
📖 4 min read☕ Coffee break read

Original authors: Sheng Li, Jingcheng Huang, Min Li

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). ✨ This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine a robot trying to have a conversation with you. If the robot's mouth moves like a stiff puppet—jerking from one shape to another without any flow—it feels creepy and unnatural. This is the "Uncanny Valley" effect: the closer a robot looks to a human, the more we notice when it doesn't move like one.

This paper is about teaching a robot how to talk with its lips in a way that feels truly human, specifically for the Chinese language. Here is the breakdown of their solution, using some everyday analogies.

The Problem: The "Stop-and-Go" Robot

Most robots today treat talking like a slideshow. They have a list of mouth shapes (like "O" for "Oh" or "M" for "Mom"). When the robot hears a sound, it just snaps to that picture.

  • The Flaw: Real humans don't snap. We slide. When we say "Moo," our lips don't just jump to an "O" shape; they round up while we are still making the "M" sound. This is called coarticulation (sounds blending together).
  • The Result: Without this blending, robot speech sounds choppy, and the lips look like they are vibrating or glitching.

The Solution: A Three-Step Recipe for Human-Like Lips

The authors built a system with three main ingredients to fix this:

1. The "Movie Script" Library (3D Dynamic Visemes)

Instead of a static photo of a mouth, the researchers created a 3D movie library for every sound in Chinese.

  • The Analogy: Imagine a traditional dictionary has a picture of a smile. This new library is like a high-definition video clip of a smile happening. It records not just the start and end, but the exact path the lips take to get there.
  • Why it matters: They mapped these movements to the ARKit standard (the same tech Apple uses for Face ID), turning complex muscle movements into 27 digital "knobs" that control the face. They then simplified the hundreds of Chinese sounds down to 14 core "lip characters" (visemes) that cover all the necessary movements.

2. The "Smoothie Blender" (Coarticulation Modeling)

This is the secret sauce. When you say a word like "Gua" (Watermelon), it's actually three sounds blended together: G-u-a.

  • The Analogy: If you just mixed ingredients in a blender one by one, you'd get chunks. You need to blend them while they are spinning.
  • How it works: The system uses a special math formula to "cross-fade" between sounds. It knows that when you are finishing the "G" sound, your lips should already start rounding for the "u." It also adjusts the speed based on how loud the sound is (energy modulation). If you shout, the lips move faster and wider; if you whisper, they move gently. This prevents the robot's lips from getting stuck or vibrating.

3. The "Translator" (Motion Mapping)

Here is the tricky part: The robot's head doesn't have 27 independent muscles like a human. It has a mechanical skeleton with only 14 moving parts (motors and cables).

  • The Analogy: Imagine you are a conductor trying to get a 27-piece orchestra to play, but you only have 14 batons. You have to tell the batons how to move to simulate the sound of the whole orchestra.
  • How it works: The system acts as a translator. It takes the complex 27-knob digital command and figures out how to combine them to move the robot's 14 physical motors. They used a mix of computer data and human experts to "calibrate" this, ensuring the robot's silicone skin stretches naturally without tearing or looking stiff.

Did it Work? (The Results)

The team tested their robot against four other methods and real human actors.

  • Smoothness: They measured "jerkiness" (how sudden the movements were). Their robot was almost as smooth as a real human, while other methods looked like they were having a seizure.
  • Accuracy: They measured how well the robot's mouth matched the sound. Their method was significantly better at matching the rhythm and shape of human speech.

The Big Picture

This research gives robots a "voice" that isn't just about the audio, but the visual performance too. By making the lips move with the fluidity of a human, it helps people understand the robot better (especially in noisy rooms) and makes the interaction feel less scary and more friendly.

In short: They stopped the robot from "snapping" its mouth and taught it to "slide" its lips, making it a much better conversational partner.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →