← Latest papers
💻 computer science

Wan-Dancer: A Hierarchical Framework for Minute-scale Coherent Music-to-Dance Generation

The paper introduces Wan-Dancer, a novel hierarchical framework that overcomes the temporal limitations of existing diffusion models to generate minute-scale, high-definition (720p/30fps), and rhythmically coherent dance videos directly from music by decoupling global keyframe planning from local temporal refinement.

Original authors: Mingyang Huang, Peng Zhang, Li Hu, Guangyuan Wang, Bang Zhang

Published 2026-07-13
📖 5 min read🧠 Deep dive

Original authors: Mingyang Huang, Peng Zhang, Li Hu, Guangyuan Wang, Bang Zhang

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine trying to choreograph a dance for a whole song, but your brain can only remember the moves for about 15 seconds before it starts to glitch, forget the dancer's face, or make them trip over their own feet. That's the current state of most AI dance generators. They are great at short bursts, but as soon as you ask for a minute-long performance, the video starts to wobble, the dancer's identity flickers, and the rhythm falls apart.

Enter Wan-Dancer, a new framework from Alibaba's Tongyi Lab that acts like a master choreographer with a superpower: it can plan the entire dance from start to finish without losing its cool.

The Problem: The "Amnesia" of Short Memories

Current AI models are like dancers with very short-term memory. If you ask them to dance for 20 seconds, they look great. But if you push them to 60 seconds or more, they suffer from "temporal drift." This means the dancer might start the song looking like a pop star and end it looking like a completely different person, or their movements might become repetitive and robotic. Previous methods tried to fix this by stitching together short clips, but that's like taping together a movie reel; you can always see the seams, and the story falls apart.

The Solution: The "Global Planner" and the "Local Dancer"

Wan-Dancer solves this by splitting the job into two distinct roles, working together in a hierarchical team:

  1. The Global Planner (The Big Picture): First, the AI looks at the entire song at once. It doesn't just guess the next move; it plans the "keyframes"—the major poses and structural moments of the dance—spread out over the whole minute. Think of this as an architect drawing the blueprint of a skyscraper before laying a single brick. This ensures the dancer knows where they are going to be at the 50-second mark, even while they are dancing at the 5-second mark.
  2. The Local Refiner (The Details): Once the blueprint is set, a second part of the system fills in the gaps. It takes those key poses and generates the smooth, high-definition frames in between. This is where the magic of fluid motion happens, ensuring the dancer doesn't teleport from one pose to another but glides through the space.

The Secret Sauce: Three Cool Tricks

To make this work, the researchers added three specific tools to their toolkit:

  • The Time-Traveling Clock (Dynamic Frame Rate): Music isn't always the same speed. Sometimes it's a slow ballad; sometimes it's a fast techno track. Wan-Dancer uses a special "time-mapped" clock (called RoPE embeddings) that tells the AI exactly how much time has passed, regardless of how many frames per second are being generated. This allows the AI to sync perfectly with a 3-second beat or a 10-second slow-motion stretch without getting confused.
  • The Motion Blur Shield (Optical Flow Loss): When a dancer spins fast, things get blurry. Old AI models would just smear the image, turning a hand into a blob. Wan-Dancer uses an "optical flow" loss function. Imagine this as a strict teacher who checks the video frame-by-frame to make sure the movement lines up perfectly. If the AI tries to blur the details during a fast spin, the teacher says, "Nope, fix that hand," ensuring the dancer stays sharp even at high speeds.
  • The Speed Dial (Motion Control): The system was trained to handle slow, medium, and fast movements differently. It learned that "medium" speed is usually the sweet spot for human dance, preventing the AI from making moves that are too slow (boring) or too fast (physically impossible and glitchy).

The Results: A Minute of Magic

The team tested this on a massive dataset of about 200 hours of high-quality dance videos, covering five distinct styles: Chinese Classical, K-Pop, Latin, Tap, and Street dance.

The results are impressive. Wan-Dancer can generate stable, 720p resolution videos at 30 frames per second that last over one minute (specifically, they demonstrated up to 160 seconds).

  • Identity: The dancer looks like the same person from the first second to the last.
  • Rhythm: The moves hit the beats of the music perfectly.
  • Versatility: It can switch between dance styles based on a simple text prompt, like "A dancer is performing a Latin dance."

They also showed that you can teach the AI a specific routine using just 16 reference videos and a technique called LoRA (Low-Rank Adaptation). This means you can make the AI mimic a specific choreography without needing to retrain the whole system from scratch.

What It's NOT (And What's Next)

It's important to note what Wan-Dancer doesn't do yet. It doesn't generate 3D skeletons that you have to render later; it goes straight to the video. However, the authors admit it's not perfect.

  • Face Consistency: While the body stays consistent, the face might still flicker slightly over very long sequences. They plan to fix this in the future.
  • Group Dancing: Right now, it's a solo act. They want to teach it how to choreograph groups of dancers interacting with each other.
  • Ethics: The team is very clear: this is for creating art, not for making fake news or deepfakes. They promise that all outputs will be clearly labeled as AI-generated.

In short, Wan-Dancer suggests that by breaking a big problem into a "big picture" plan and a "fine detail" execution, we can finally make AI dance to the whole song without losing the beat. It's a significant step forward, moving us from 15-second clips to full-minute performances that actually look like a real human is dancing.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →