← Latest papers
💻 computer science

MotionWeaver: Holistic 4D-Anchored Framework for Multi-Humanoid Image Animation

MotionWeaver is an end-to-end framework that overcomes the limitations of existing single-human image animation methods by introducing unified motion representations and a holistic 4D-anchored paradigm to achieve state-of-the-art results in complex multi-humanoid scenarios involving diverse forms, interactions, and occlusions.

Original authors: Xirui Hu, Yanbo Ding, Jiahao Wang, Tingting Shi, Yali Wang, Guo Zhi Zhi, Weizhan Zhang

Published 2026-02-17
📖 4 min read☕ Coffee break read

Original authors: Xirui Hu, Yanbo Ding, Jiahao Wang, Tingting Shi, Yali Wang, Guo Zhi Zhi, Weizhan Zhang

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are a director trying to film a movie with a cast of characters. Some are humans, some are robots, and some are cartoon animals. You have a photo of each character and a video of how you want them to move. Your goal is to make the photos come alive and dance, fight, or hug exactly like the video, even when they bump into each other or hide behind one another.

For a long time, computers were great at animating one person. But as soon as you added a second person, or a robot, or an alien, the computer got confused. It would mix up who was who, or it would make them walk through each other like ghosts.

Enter MotionWeaver. Think of it as a master puppeteer who can handle an entire troupe of different characters at once without getting tangled up. Here is how it works, broken down into simple concepts:

1. The Problem: The "Identity Crisis"

Imagine you have two dancers, Alice and Bob. In old computer programs, if you told the computer "Alice, do Bob's dance," the computer might get confused because it was looking at the shape of the body to decide who was who.

  • The Old Way: It was like trying to recognize a friend only by their height. If a tall person and a short person swapped places, the computer got lost.
  • The New Way (MotionWeaver): It separates the dance steps from the dancer. It realizes, "Oh, these are the steps, and that is Alice's face. I will give Alice's face these steps." It doesn't care if Alice is a human, a robot, or a dragon; it just knows how to move the specific character.

2. The Secret Sauce: The "4D Map"

Most video generators work in 2D (flat images) and time. They don't really understand depth.

  • The Analogy: Imagine a 2D drawing of two people hugging. From the front, you can't tell who is in front and who is in back. A 2D computer might draw them melting into a blob.
  • MotionWeaver's Trick: It builds a 4D Map (3D space + time). It understands that "Person A is behind Person B."
    • It uses a special "Depth Sensor" (called the Hyper-Scene Integrator) that acts like a pair of 3D glasses for the computer.
    • When two characters cross paths, the computer knows exactly who should be visible and who should be hidden, just like in real life.

3. The Training: "The Dance School"

To teach this computer to be so good, the authors didn't just use YouTube videos of people dancing. They built a massive, custom dance school called MultiHuman46.

  • The Dataset: They collected 46 hours of videos showing people (and AI-generated characters) interacting, fighting, and dancing together.
  • The Benchmark: They created a test called DualDynamics with 300 videos of two characters interacting. It's like a final exam where the computer has to prove it can handle complex moves without the characters glitching or swapping identities.

4. The Learning Process: "The Two-Stage Lesson"

The computer learns in two distinct phases, similar to how a human learns to draw:

  1. The Big Picture (High Noise): At the start of learning, the computer focuses on the big shapes and the order of things. "Who is standing where? Who is blocking whom?" It learns the occlusion (who is in front) first.
  2. The Details (Low Noise): Later, it focuses on the fine details of the movement. "How does the arm bend? How does the foot land?"

Why This Matters

Before MotionWeaver, if you wanted to animate a scene with a robot fighting a knight, you had to do it frame-by-frame by hand, which takes forever.

  • MotionWeaver allows you to upload a photo of a robot and a photo of a knight, show it a video of a fight, and it instantly generates a video where the robot and knight fight realistically, respecting who is behind whom, without the robot turning into a human or the knight disappearing.

In Summary:
MotionWeaver is like a universal translator for movement. It takes the "language" of motion (how things move in 3D space) and teaches it to any character, whether it's a human, a robot, or a cartoon monster, ensuring they never lose their identity and always know who is standing in front of whom. It turns the chaotic mess of multi-character animation into a smooth, coordinated performance.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →