Mitigating Error Accumulation in Co-Speech Motion Generation via Global Rotation Diffusion and Multi-Level Constraints
This paper introduces GlobalDiff, a novel diffusion-based framework that mitigates error accumulation in long-horizon co-speech gesture generation by operating directly in global joint rotation space and employing a multi-level constraint scheme to ensure structural integrity and temporal consistency.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine a virtual character speaking to you. For the conversation to feel real, the character cannot just move its mouth; the entire body must participate. The shoulders shift, the hands gesture, and the posture changes to match the rhythm and emotion of the voice. This is the goal of co-speech motion generation: teaching computers to create full-body movements that naturally sync with spoken words. For years, researchers have tried to solve this by building digital skeletons, where the movement of a hand is calculated by adding up the rotations of the shoulder, the elbow, and the wrist in a chain. While this approach mimics how bones are connected, it carries a hidden flaw. Just as a small mistake in a long line of people passing a message can result in a completely wrong final message, tiny errors in calculating the rotation of a shoulder can grow larger and larger as they travel down the arm. By the time the computer calculates the position of the fingertips, the error can be so large that the hand looks twisted, broken, or floating in impossible ways.
A team of researchers at Alibaba Group and Nanyang Technological University has proposed a new way to solve this problem, moving away from the chain-reaction method entirely. Instead of calculating how each joint turns relative to the one before it, their new system, called GlobalDiff, decides the final direction of every single joint in the room at the same time. Think of it as telling every joint exactly where to point in the world, rather than telling the elbow where to point relative to the shoulder. This approach stops the small errors from piling up, ensuring that the hands and feet stay stable and accurate even during long, expressive conversations. However, giving every joint total freedom creates a new risk: the character might twist its limbs into shapes that no human body could physically achieve. To fix this, the researchers added a set of invisible rules that act like a guide, checking that the bones stay the right length, the angles between limbs make sense, and the movement flows smoothly over time.
The researchers tested this new system on a large collection of recorded conversations involving over 1,700 dialogue sequences. They found that by predicting the global direction of joints directly, the computer could generate movements that were far more stable than previous methods. In their tests, the new system reduced errors in the final motion by 46 percent compared to the best existing tools. When they looked closely at the results, the difference was clear. In older systems, the fingers often looked like they were flipping inside out or the hands would drift away from the body. With the new method, the hands remained in natural positions, and the gestures matched the speaker's voice with a precision that felt human. The system also learned to handle different speakers, maintaining this high quality whether the voice belonged to a man or a woman, a native speaker or someone with a different accent.
To ensure the movements were not just stable but also physically possible, the team introduced three layers of checks. First, they placed virtual markers around each joint to make sure the rotation was precise, preventing the limbs from twisting in unnatural ways. Second, they measured the angles between all the bones in the body to ensure the skeleton remained intact, preventing the character from bending its spine or limbs in impossible directions. Third, they analyzed the timing of the movement to ensure the gestures flowed with the rhythm of the speech, avoiding jerky or erratic motions. When they removed these checks one by one, the quality of the movement dropped significantly, proving that all three layers were necessary to create a convincing result.
The study demonstrates that it is possible to generate realistic, long-duration body language for virtual characters without the accumulation of errors that has plagued the field for so long. By treating the entire body as a collection of independent parts that are guided by strict physical rules, the researchers have created a system that produces smooth, anatomically correct, and expressive motion. This work suggests that the future of virtual avatars may not rely on complex chains of calculations that are prone to failure, but rather on a direct, global understanding of how every part of the body should move in space. The result is a digital presence that can speak and move with a naturalness that brings us closer to the seamless interaction between humans and machines.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.