Coordinating Multiple Conditions for Trajectory-Controlled Human Motion Generation
The paper proposes CMC, a decoupled two-stage framework that coordinates text and trajectory conditions for human motion generation by first ensuring accurate trajectory following through a simplified representation and then completing full-body motion via a text-conditioned inpainting model enhanced with a Selective Inpainting Mechanism to achieve state-of-the-art performance.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are trying to direct a movie scene where an actor must perform a specific action (like "dancing") while walking along a very specific path drawn on the floor (like a triangle).
The Problem: The "Two-Headed" Director
Previous methods tried to do this by giving the actor two directors shouting at the same time. One director yells, "Dance!" (the text), and the other yells, "Stay on the line!" (the trajectory). Because these instructions are so different—one is a vague feeling, the other is a precise geometric line—the actor gets confused. They might start dancing but drift off the line, or they might stay perfectly on the line but stop dancing entirely.
Furthermore, these old methods tried to control every tiny detail of the actor's body (their speed, their rotation, their foot position) all at once. This was like trying to steer a car by adjusting the wheels, the engine, and the steering wheel simultaneously while driving; it made the movement wobbly and unstable.
The Solution: CMC (The "Divide and Conquer" Strategy)
The authors propose a new system called CMC (Coordinating Multiple Conditions). Instead of one confused director, they use a two-step assembly line that separates the tasks.
Step 1: The "Skeleton" Architect (Trajectory Control)
First, the system ignores the fancy details of the dance. It focuses only on the "skeleton" of the movement: the hips and the specific joints you want to control (like the hands or feet).
- The Analogy: Imagine an architect drawing a simple wireframe of a person walking along a tightrope. They don't worry about the person's shirt color or facial expression yet; they just ensure the wireframe stays perfectly on the tightrope.
- The Trick: They use a "simplified" version of the body data. Instead of tracking speed and rotation, they only track where the joints are. This makes the path-following incredibly stable and accurate.
Step 2: The "Costume Designer" (Motion Completion)
Once the wireframe is locked in place on the tightrope, the system moves to the second stage. Now, it brings in the "text" instructions (e.g., "dancing").
- The Analogy: The costume designer takes the fixed wireframe and fills in the rest of the body. They add the arms, the legs, and the style of movement based on the text description.
- The Safety Net: Because the wireframe is already frozen in the right spot, the costume designer doesn't have to worry about the person drifting off the tightrope. They can focus entirely on making the dance look natural and expressive.
The Secret Sauce: "Selective Inpainting"
There was a risk that the "Costume Designer" (Step 2) would get too good at just copying the wireframe and forget how to dance on their own. To fix this, the authors introduced a training technique called SIM (Selective Inpainting Mechanism).
- The Analogy: Imagine teaching a student to paint a portrait. If you only let them paint over a pre-drawn sketch, they might forget how to draw a face from scratch. So, the teacher randomly tells them: "Today, paint over the sketch (inpainting)." "Tomorrow, draw a whole face from scratch without a sketch (text-to-motion)."
- The Result: This keeps the model flexible. It learns to fill in the blanks when given a path, but it also remembers how to create natural movements from scratch, preventing it from becoming rigid or "overfit."
The Outcome
By separating the "path" from the "style," CMC solves the conflict. The result is a human motion that follows the path perfectly (like a tightrope walker) while still looking like a natural, expressive dance (like a performer).
The paper shows that this method works better than previous "two-headed director" approaches on standard datasets, creating movements that are both accurate to the path and realistic in their appearance.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.