← Latest papers
💻 computer science

ONE-SHOT: Compositional Human-Environment Video Synthesis via Spatial-Decoupled Motion Injection and Hybrid Context Integration

ONE-SHOT is a parameter-efficient framework that achieves fine-grained, compositional human-environment video synthesis by decoupling human dynamics from environmental cues through spatial-decoupled motion injection and hybrid context integration, eliminating the need for heavy 3D pre-processing while ensuring long-horizon consistency.

Original authors: Fengyuan Yang, Luying Huang, Jiazhi Guan, Quanwei Yang, Dongwei Pan, Jianglin Fu, Haocheng Feng, Wei He, Kaisiyuan Wang, Hang Zhou, Angela Yao

Published 2026-04-02
📖 5 min read🧠 Deep dive

Original authors: Fengyuan Yang, Luying Huang, Jiazhi Guan, Quanwei Yang, Dongwei Pan, Jianglin Fu, Haocheng Feng, Wei He, Kaisiyuan Wang, Hang Zhou, Angela Yao

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you want to make a movie where a specific person (let's say, your friend holding a sword) performs a Tai Chi routine inside a specific location (like a museum), but you want to be able to swap the person, change the sword, or move the camera around without the whole scene falling apart.

Making this kind of video with AI has been like trying to bake a cake where the flour, eggs, and oven temperature are all glued together. If you try to change the flavor (the person), the whole cake collapses. If you try to change the oven (the background), the eggs scramble.

ONE-SHOT is a new "smart kitchen" that un-glues these ingredients so you can mix and match them perfectly. Here is how it works, using simple analogies:

1. The Problem: The "Over-Conditioning" Mess

Previous AI video tools were like a strict chef who demanded you give them a perfectly pre-assembled 3D model of the person and the room before they could start cooking.

  • The Issue: It was hard to get that perfect 3D model (boring, technical work). Also, because the chef was so focused on the 3D model, they forgot how to be creative. If you asked for a "sleek robot dog," the AI might just give you a blurry dog because it was too busy trying to fit the 3D math.

2. The Solution: The "Decoupled" Kitchen

ONE-SHOT changes the recipe. Instead of gluing everything together, it separates the ingredients into three distinct bowls:

  1. The Person's Dance (Motion): How they move.
  2. The Room (Environment): Where they are.
  3. The Script (Text): What they are doing.

The AI learns to look at these bowls separately and then mix them together at the very last second. This means you can swap the "Dance" bowl with a different dance, or the "Room" bowl with a different room, and the AI knows exactly how to blend them without getting confused.

3. The Secret Sauce: Three Magic Tools

A. The "Ghost Stage" (Canonical-Space Injection)

Imagine you want to put a dancer on a stage.

  • Old Way: You try to build the stage around the dancer's specific body shape. If the dancer is tall, the stage gets weirdly tall.
  • ONE-SHOT Way: The AI creates a "Ghost Stage" (a standard, neutral space) where the dancer's moves are recorded first. Then, it projects that performance onto your actual museum background.
  • Why it's cool: The dancer doesn't need to know what the museum looks like. The AI just says, "Okay, the dance happens here in the museum," and it fits perfectly. This stops the AI from getting "over-conditioned" (confused by too many rules).

B. The "Smart Ruler" (Dynamic-Grounded-RoPE)

This is the paper's most technical part, but think of it as a magic ruler that stretches and shrinks.

  • The Problem: The "Ghost Stage" is a small, standard size. The "Museum" is huge and complex. If you try to paste the small dance onto the big museum, the feet might end up floating in the air or the head might be in the ceiling.
  • The Fix: ONE-SHOT uses a "Smart Ruler" that looks at where you want the person to stand (a simple box you draw). It then stretches the "Ghost Stage" coordinates to match that box perfectly.
  • Result: The person's feet stay on the floor, and their hands stay in the air, no matter how big or small the room is. No complex 3D math required beforehand!

C. The "Memory Bank" (Hybrid Context Integration)

If you ask an AI to make a 1-minute video, it often forgets what the person looked like at the 10-second mark. By the end, the person might look like a different human entirely.

  • The Fix: ONE-SHOT has a Memory Bank.
    • Static Memory: It keeps a photo of the person's face and body to remember "Who is this?"
    • Dynamic Memory: It keeps a running note of "What did the room look like 5 seconds ago?"
  • Result: You can generate a video that is minutes long, and the person will look exactly the same from start to finish, and the museum won't suddenly turn into a forest.

4. What Can You Do With This?

Because the ingredients are separated, you can do "Compositional" magic:

  • Swap the Actor: Put a different person in the same Tai Chi video.
  • Swap the Move: Make the same person do a dance instead of Tai Chi.
  • Swap the Scene: Put the Tai Chi master in a forest, a spaceship, or a museum.
  • Text Magic: Just type "Make the sword glow," and the AI adds the glow without breaking the video.

The Bottom Line

ONE-SHOT is like a Lego set for video. Instead of trying to sculpt a whole statue out of clay (which is hard and rigid), it gives you pre-made, flexible blocks for people, places, and actions. You can snap them together in any combination, and the result is a high-quality, realistic video that stays consistent, even if you keep building for a long time.

It solves the three biggest headaches of current AI video:

  1. Too much 3D work: It skips the heavy 3D modeling.
  2. Too rigid: It lets the AI be creative again.
  3. Short attention span: It can make long videos without the characters turning into aliens.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →