← Latest papers
💻 computer science

Preserve, Reveal, Expand: Faithful 4D Video Editing with Region-Aware Conditioning

The paper introduces PREX, a region-aware framework that addresses the evidence-role mismatch in 4D video editing by decomposing spatiotemporal volumes into Preserve, Reveal, and Expand roles to faithfully synthesize disoccluded content while maintaining source consistency, alongside the PREBench benchmark for evaluation.

Original authors: Zhangchi Hu, Wenzhang Sun, Xiangchen Yin, Jiahui Yuan, Chunfeng Wang, Hao Li, Kun Zhan, Xiaoyan Sun

Published 2026-05-21
📖 5 min read🧠 Deep dive

Original authors: Zhangchi Hu, Wenzhang Sun, Xiangchen Yin, Jiahui Yuan, Chunfeng Wang, Hao Li, Kun Zhan, Xiaoyan Sun

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have a home video of your family at a park. Now, imagine you want to edit that video so that a tree in the background moves to the left, or so that the camera zooms in to see a bird that was previously hidden behind a bush.

In the world of AI video editing, this is called 4D Video Editing. It's like taking a flat movie and turning it into a 3D world you can walk around inside. But there's a big problem: when AI tries to do this, it often gets confused about what parts of the video are "real" (from the original recording) and what parts are "new" (imagined by the AI).

This paper introduces a new method called PREX (which stands for Preserve, Reveal, Expand) to fix this confusion. Here is how it works, using simple analogies:

The Problem: The "Confused Chef"

Imagine a chef (the AI) trying to cook a meal based on a recipe (the video edit).

  • The Mistake: The chef is given a single bowl of ingredients that mixes fresh, real vegetables (from the original video) with blurry, fake plastic vegetables (generated by the computer).
  • The Result: The chef tries to use all of them at once. The real vegetables get mushy and lose their taste (the original video gets distorted), and the fake vegetables look weird and ghostly (the new parts look wrong).

The paper calls this the "Evidence-Role Mismatch." The AI doesn't know which parts of the video to trust and which parts to invent.

The Solution: The "Three-Zone Kitchen"

PREX solves this by telling the AI to stop treating the whole video as one big bowl. Instead, it divides the video into three distinct zones, like a kitchen with three separate stations:

  1. The Preserve Zone (The "Do Not Touch" Station):

    • What it is: These are the parts of the video that are still visible and unchanged (like your family member who didn't move).
    • The Rule: The AI must copy these pixels exactly from the original video. No guessing, no changing. It's like a strict "Do Not Touch" sign on a museum exhibit.
    • Goal: Keep the original look and feel perfectly intact.
  2. The Reveal Zone (The "Hidden Treasure" Station):

    • What it is: These are parts of the scene that were hidden before but are now visible because something moved (like the bird behind the bush).
    • The Rule: The AI knows these areas exist in the 3D world, but it has no photo of them. It must "paint" them in using clues from the surrounding area.
    • Goal: Fill in the blanks so they look like they belong, without copying the wrong things (like the tree that moved away).
  3. The Expand Zone (The "New Horizon" Station):

    • What it is: These are parts of the video that appear because the camera moved to a new angle, showing areas that were never in the original video at all (like the sky above the tree).
    • The Rule: The AI has to imagine this new world from scratch, making sure it fits the style and physics of the rest of the video.
    • Goal: Create something new that doesn't look like a glitch or a copy-paste error.

How PREX Works

PREX acts like a smart foreman for the AI chef.

  • It creates a map that tells the AI exactly which pixels belong to which zone (Preserve, Reveal, or Expand).
  • It gives the AI a confidence score for every pixel. If the AI is 100% sure a pixel is from the original video, it locks it in place. If the AI is unsure, it knows it's allowed to invent that part.
  • It uses a special adapter (a small tool added to the AI) to make sure the "Do Not Touch" areas stay perfect while the "New" areas are created smoothly.

The Proof: PREBench

To make sure this actually works, the authors built a new testing ground called PREBench.

  • Instead of just asking, "Does this video look good overall?" (which is vague), PREBench asks specific questions:
    • "Did the original family member stay looking exactly the same?" (Preserve check).
    • "Is the bird behind the bush looking real, or is it a ghost?" (Reveal check).
    • "Does the new sky look stable, or is it flickering?" (Expand check).

The Results

When they tested PREX against other AI video editors:

  • Less Ghosting: The "ghosts" (weird, semi-transparent artifacts) in the newly revealed areas disappeared.
  • Better Preservation: The parts of the video that shouldn't change stayed exactly the same, rather than getting blurry or distorted.
  • Smoother Expansion: The new parts of the video looked more natural and didn't look like someone just stretched the old picture.

In short, PREX teaches the AI to know the difference between what it knows (the original video) and what it needs to imagine (the new parts), resulting in a much cleaner, more realistic video edit.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →