← Latest papers
🤖 AI

Robust Motion Generation using Part-level Reliable Data from Videos

This paper addresses the challenge of occlusions and off-screen captures in video-based motion data by proposing a robust part-aware masked autoregression model that leverages only credible part-level information to generate high-quality human motion, accompanied by the introduction of a new large-scale benchmark called K700-M.

Original authors: Boyuan Li, Sipeng Zheng, Bin Cao, Ruihua Song, Zongqing Lu

Published 2026-08-26
📖 6 min read🧠 Deep dive

Original authors: Boyuan Li, Sipeng Zheng, Bin Cao, Ruihua Song, Zongqing Lu

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine trying to teach a computer to move a digital human just by reading a sentence. You might type, "The person is dancing," and expect the screen to show a fluid, realistic performance. This field, known as text-to-motion generation, has made remarkable progress, but it has hit a wall. The best results so far come from training computers on special, high-quality recordings of real people moving, captured with expensive sensors in controlled studios. While these recordings are perfect, they are rare, limited in variety, and often miss the messy, complex ways people actually move in the real world. To break through this barrier, researchers have turned to the internet, hoping to learn from the billions of ordinary videos available online. The idea is simple: if a computer can learn from the vast library of the web, it could understand a much wider range of actions. But there is a catch. When a computer tries to figure out how a person is moving in a regular video, it often cannot see the whole body. A person might be partially hidden by a tree, cut off by the edge of the camera frame, or blocked by another person. In these moments, the computer has to guess where the missing limbs are.

This guessing creates a fundamental problem for learning. If a computer is shown a video where only the upper body is visible, and it is forced to learn from a full-body guess that includes invisible legs, it might learn the wrong things. It could start believing that the invisible legs are real evidence, rather than just a guess. For a long time, researchers faced a difficult choice: either throw away every video clip that wasn't perfectly clear, losing most of the data, or keep all the clips and hope the computer could ignore the bad guesses. Both approaches had serious flaws. Throwing away clips meant losing valuable information about specific movements, like how a person moves their arms while sitting. Keeping everything meant teaching the computer with false information. A team of researchers has now found a way to solve this by changing how the computer looks at the data. Instead of treating the whole video clip as either good or bad, they taught the computer to look at the body part by part.

The researchers developed a new system that acts like a careful editor. When the system analyzes a video, it first checks which parts of the body are actually visible and which parts are hidden. If a person's legs are blocked by a table, the system knows to trust the movement of the arms and torso but to ignore the computer's guess about the legs. It treats the visible parts as solid facts and the hidden parts as blanks to be filled in later. This approach allows the system to learn from thousands of videos that were previously considered too messy to use. By focusing only on the parts of the body that the camera actually sees, the system builds a reliable understanding of movement without being confused by its own guesses. It is like learning to recognize a bird by studying only the wings when the tail is hidden, rather than trying to memorize a drawing of the whole bird that includes a made-up tail.

To test this idea, the team created a massive new collection of data from the internet, containing nearly two hundred thousand pairs of video clips and text descriptions. They analyzed these clips and found that in more than three-quarters of the frames, at least one part of the body was not fully visible. This confirmed that the problem of missing information is not rare; it is the norm for web videos. Using this dataset, they trained their new system and compared it against older methods that either threw away the messy clips or tried to learn from all of them without filtering. The results were clear. The new system, which pays attention to which parts are visible, produced much more accurate and realistic movements than the older methods. It was able to generate motions that were not only more faithful to the text description but also smoother and more natural.

The study showed that keeping the messy clips and teaching the computer to ignore the invisible parts was far better than throwing the clips away. When the researchers tested their system on a large pool of two hundred thousand clips, it significantly outperformed the best existing methods. The older methods that kept all the data made many mistakes because they tried to learn from the invisible parts, while the methods that threw away the bad data lost too much variety. The new system managed to keep the variety of the large dataset while avoiding the mistakes. It learned that a person sitting and playing drums might have their legs hidden, but their arm movements are still real and useful to learn from. By focusing on the reliable parts, the system could piece together a complete picture of the movement without being misled by the missing pieces.

This work suggests a new way to teach computers about the physical world using imperfect data. It shows that we do not need perfect, studio-quality recordings to teach a machine how to move. Instead, we can use the vast, messy collection of videos from the internet if we teach the machine to be smart about what it trusts. The system does not need to see the whole body to learn how a person moves; it just needs to see enough of the body to understand the action. This approach opens the door to training much larger and more capable systems in the future, capable of understanding a wider range of human behaviors than ever before. The researchers found that by respecting the limits of what the camera can see, the computer can learn more, not less. This shift in strategy, from demanding perfect data to learning from partial evidence, marks a significant step forward in teaching machines to understand human motion.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →