← Latest papers
🤖 AI

GroundShot: Visually Consistent Multi-Shot Long Video Generation via Entity-Grounded Shot Scheduling

GroundShot is a training-free, model-agnostic framework that enhances visually consistent multi-shot long video generation by dynamically scheduling shot order and maintaining an online entity-level visual memory to ground and verify entity references, thereby preventing visual drift across shots.

Original authors: Yixuan Lai, Tianjia Shao, Kun Zhou, Weijia Dou, Siyu Zhu, Jingdong Wang

Published 2026-06-23
📖 5 min read🧠 Deep dive

Original authors: Yixuan Lai, Tianjia Shao, Kun Zhou, Weijia Dou, Siyu Zhu, Jingdong Wang

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are directing a movie, but instead of hiring actors and a camera crew, you are asking a magical AI to draw every single scene based on a script you wrote.

The biggest problem with current AI movie-makers is memory loss. If you ask the AI to draw a scene with a "red-haired woman in a blue coat" in Scene 1, and then ask for her again in Scene 10, the AI often forgets what she looked like. Maybe her hair turns brown, the coat becomes green, or she suddenly has a different face. This happens because the AI tries to remember the last thing it drew, and small mistakes pile up like a game of "telephone," until the character is unrecognizable.

GroundShot is a new system designed to fix this. It doesn't just let the AI draw scenes in order; it acts like a smart director and a meticulous archivist working together.

Here is how it works, using simple analogies:

1. The "Best Photo" Rule (Quality-Aware Scheduling)

In a normal movie, you shoot scenes in the order they happen in the story. But GroundShot realizes that not all shots are good for remembering what a character looks like.

  • The Problem: If the first time a character appears in the story is a tiny, blurry figure in the distance (a "wide shot"), that's a terrible photo to remember them by. If the AI uses that blurry photo as a reference for the rest of the movie, the character will stay blurry or distorted.
  • The GroundShot Solution: The system looks at the whole script first. It says, "Wait, let's not shoot the story in order yet. Let's shoot the close-up of the character's face first, even if that scene happens later in the story."
  • The Analogy: Imagine you are trying to teach someone to recognize your dog. You wouldn't start by showing them a photo of the dog from 100 yards away in the rain. You would first show them a clear, high-quality photo of the dog's face. GroundShot generates that "clear face photo" first, locks it in as the Master Reference, and then uses that perfect image to guide the drawing of every other scene, no matter where it appears in the story.

2. The "Specialized Filing Cabinet" (Entity-Level Memory)

Older systems try to remember the entire picture of a scene. This is like trying to remember a whole room to find one specific chair. It's messy and confusing.

  • The GroundShot Solution: GroundShot breaks the scene down into individual items: the Character, the Object (like a jacket), and the Location (like a restaurant).
  • The Analogy: Instead of one giant photo album, GroundShot has three separate filing cabinets:
    • Cabinet A (Characters): Holds the best photo of the "Red-Haired Woman."
    • Cabinet B (Objects): Holds the best photo of the "Blue Jacket."
    • Cabinet C (Locations): Holds the best photo of the "Restaurant."
      When it needs to draw a new scene, it pulls the specific "Red-Haired Woman" file and the "Blue Jacket" file. It doesn't get confused by the background or other people in the room.

3. The "Quality Control Inspector" (Verification & Grounding)

Just because the AI draws something doesn't mean it's good enough to be a reference.

  • The Process: After the AI draws a scene, GroundShot acts like a strict film editor. It checks: "Is the face clear? Is the coat the right color? Did the character disappear?"
  • The Filter: If the drawing is blurry, cut off, or weird, GroundShot throws it in the trash. It only keeps the perfect versions to put into the filing cabinets. If a character is only seen from the back, GroundShot knows that's a bad reference for their face, so it waits until it can generate a front-facing shot before updating the memory.

4. The "Dynamic Library" (Smart Updates)

Sometimes, a character needs to look different (e.g., smiling vs. frowning, or seen from the side).

  • The Solution: GroundShot keeps the "Master Reference" (the perfect face) safe and unchanged. But, if a new scene requires a side view, it adds a supplementary file to the cabinet that shows the side view, as long as it still matches the Master Reference.
  • The Analogy: Think of it like a 3D model. You have the main blueprint (the Master Reference). If you need to build a wall on the side, you add a side-view blueprint, but you never change the main blueprint. This ensures the character never "drifts" into looking like a different person.

The Result

By doing all this, GroundShot creates long videos where characters, clothes, and places stay exactly the same from the first second to the last. It doesn't need to be retrained or taught new tricks; it just organizes the process smarter.

In short: GroundShot stops the AI from playing "telephone" with its own memories. Instead, it creates a "Master Reference" for every character and object, generates the best possible version of them first, and then uses that perfect version to ensure consistency throughout the entire movie.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →