← Latest papers
💻 computer science

ReplicateAnyScene: Zero-Shot Video-to-3D Composition via Textual-Visual-Spatial Alignment

The paper introduces ReplicateAnyScene, a zero-shot framework that transforms casually captured videos into compositional 3D scenes by aligning textual, visual, and spatial priors from foundation models, while also proposing the C3DR benchmark to evaluate reconstruction quality.

Original authors: Mingyu Dong, Chong Xia, Mingyuan Jia, Weichen Lyu, Long Xu, Zheng Zhu, Yueqi Duan

Published 2026-04-15
📖 5 min read🧠 Deep dive

Original authors: Mingyu Dong, Chong Xia, Mingyuan Jia, Weichen Lyu, Long Xu, Zheng Zhu, Yueqi Duan

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you walk into a messy room, glance at it for a few seconds, and then instantly close your eyes. In your mind, you don't just see a blurry mess; you see a chair, a table, a bag, and a lamp. You know exactly where the chair is sitting, how heavy it feels, and that the lamp is resting on the table, not floating in mid-air. You could even build a perfect 3D model of that room in your head, piece by piece.

For a long time, computers have been terrible at this. They could either take a photo and make a flat 2D picture look 3D (like a hologram), or they could build a 3D model if you gave them a million photos and a lot of human help. But they couldn't just "watch" a video, understand what the objects are, and build a realistic, separate 3D model of every single item in the scene automatically.

Enter ReplicateAnyScene. Think of this as a super-smart, fully automated digital architect that can watch a casual video of a room and instantly build a perfect, interactive 3D version of it, object by object.

Here is how it works, broken down into five simple steps using a creative analogy:

The Analogy: Building a LEGO City from a Video

Imagine you are trying to rebuild a city using only a shaky video taken by a tourist. You have a team of experts (the AI models) helping you. Here is your five-step construction crew:

1. The Scout (Progressive Object Discovery)

The Problem: If you show the whole video to a computer at once, it gets overwhelmed and forgets things.
The Solution: The computer acts like a scout. Instead of watching the whole movie, it picks the best 20 frames (like the most interesting snapshots). It asks a "smart librarian" (a Vision-Language Model): "What objects do you see here?"
The Magic: If the librarian says "sofa" in one frame and "couch" in another, the computer knows they are the same thing. It creates a clean, non-repeating shopping list of everything in the room: Chair, Table, Bag, TV.

2. The Detective (Spatial-Guided Visual Deduplication)

The Problem: In a video, a chair might be hidden behind a table for a second, then reappear. A dumb computer might think, "Oh, that's a new chair!" and create two chairs instead of one.
The Solution: The computer acts like a detective who knows that real objects don't teleport. It takes the 2D pictures of the chair and lifts them into 3D space. If two "chairs" occupy the same physical space in 3D, the detective merges them into one. It ensures that every object in the final model is unique and real.

3. The Photographer (Optimal-View 3D Asset Generation)

The Problem: To build a 3D model of a chair, you need a good photo. But if you pick a photo where the chair is tiny or blocked, the 3D model will look weird.
The Solution: The computer acts like a professional photographer. It looks at all the frames where the chair appears and asks: "Which angle shows the most of the chair's surface?" It picks that perfect shot and uses a powerful AI artist to generate a high-quality 3D model of just that chair, complete with textures and details.

4. The Glue (Iterative Visual-Spatial Alignment)

The Problem: Now you have a 3D chair and a 3D table, but they are floating in empty space. The computer needs to put them in the right spot.
The Solution: This is the "glue" stage. The computer tries to fit the 3D chair into the video. It projects the 3D chair onto the video frame and checks: "Does the edge of my 3D chair match the edge of the real chair in the video?" If it's off, it nudges the chair slightly and tries again. It does this over and over (like a puzzle solver) until the 3D model fits the video perfectly.

5. The Physicist (Semantic-Aware Scene Refinement)

The Problem: Sometimes, the computer gets the physics wrong. It might put a chair floating in the air, or a lamp hanging from the ceiling upside down.
The Solution: The computer calls in a "physicist" (another smart AI). It asks: "Does this make sense?"

  • If the chair is floating, the physicist says, "Gravity! Put it on the floor."
  • If the lamp is upside down, it says, "Lamps sit on tables, they don't hang from the ceiling."
    It makes tiny adjustments to ensure the scene looks physically real and logical.

Why is this a big deal?

Before this, building 3D worlds for robots, video games, or virtual reality required:

  • Manual labor: Humans had to draw every object.
  • Special equipment: You needed expensive depth cameras.
  • Simple scenes: It only worked on boring, empty rooms.

ReplicateAnyScene changes the game. It is Zero-Shot, meaning it doesn't need to be trained on specific rooms; it can handle any room you show it. It is Fully Automated, meaning you just press "play" on a video, and it does the rest.

The Result

The researchers also built a new "test track" called C3DR to grade these 3D models. Their system scored higher than any previous method, creating scenes that are not only visually beautiful but also physically correct.

In short: This technology turns a simple video into a fully interactive, 3D world where every object is real, placed correctly, and ready for robots or gamers to explore. It's like giving a computer the human ability to "see" and "understand" the 3D world around us.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →