RoDyGS: Robust Dynamic Gaussian Splatting for Casual Videos
This paper introduces RoDyGS, a method that reconstructs dynamic 3D scenes from casual monocular videos by explicitly separating static and dynamic elements with spatiotemporal regularization, while also proposing a new comprehensive benchmark, Kubric-MRig, to evaluate such pose-free dynamic novel view synthesis approaches.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are holding your phone and recording a casual video of a busy park. You see people walking, dogs running, and trees swaying in the wind. Now, imagine trying to turn that flat, 2D video into a 3D world where you can walk around the trees and watch the dogs from a different angle. This is incredibly hard because your video only shows one perspective, and the camera itself is moving, making it difficult to tell what is moving because the camera moved versus what is moving on its own.
This paper introduces RoDyGS, a new tool designed to solve this puzzle. Think of RoDyGS as a "digital time-traveling sculptor" that takes your shaky, handheld video and rebuilds the entire 3D scene, including how objects move over time.
Here is how it works, broken down into simple concepts:
1. The "Static vs. Dynamic" Sorting Hat
The biggest problem with these videos is confusion. Is that tree moving because the wind blew it, or because you walked past it?
- The Analogy: Imagine a chaotic party where some guests are standing still (the walls, the floor) and others are dancing wildly (people, cars). If you try to describe the whole room at once, it's a mess.
- The Solution: RoDyGS puts on a "sorting hat." It immediately separates the scene into two groups: Static (the background that doesn't move) and Dynamic (the things that do). By treating them separately, the system stops getting confused about whether the camera is moving or the object is moving.
2. The "Ghost-Busting" Regularization
When you try to reconstruct a moving object from a single video, parts of it often disappear behind other things (occlusion). If a person walks behind a car, the system loses track of them for a moment. Without help, the 3D model might glitch, turning the person into a blob or a ghostly smear.
- The Analogy: Imagine trying to guess the shape of a toy car while someone keeps covering it with a blanket. If you just guess randomly, you might think the car is melting.
- The Solution: RoDyGS uses "rules of physics" (called regularization) to keep the model honest.
- Distance-Preserving: It reminds the system, "Hey, the distance between the car's wheels shouldn't change just because you can't see them for a second." It keeps the object rigid.
- Surface Smoothing: It ensures the surface of the object stays smooth, like a real ball, rather than looking like a jagged, pixelated rock.
- Steady Motion: It assumes that objects don't teleport or jitter wildly between frames. If a ball is rolling, it keeps rolling smoothly, not vibrating in place.
3. The "Self-Correcting Camera"
Usually, to build a 3D model, you need to know exactly where the camera was at every second. But in casual videos, you don't have that data.
- The Analogy: It's like trying to draw a map of a city while walking through it blindfolded, only knowing what you see in front of you.
- The Solution: RoDyGS acts like a detective. It looks at the static parts of the scene (the ground, the buildings) to figure out where the camera must have been. It constantly corrects its own guess about the camera's path while simultaneously building the 3D model.
4. A New "Training Gym" (Kubric-MRig)
The authors realized that existing tests for these tools were too easy or unrealistic. They were like testing a race car on a flat, empty parking lot.
- The Analogy: They built a new "obstacle course" called Kubric-MRig. This is a synthetic world generated by a computer where they can control everything. They created scenes with wild camera movements, objects flying around, and complex backgrounds, all with a "perfect answer key" (ground truth) so they could prove their tool actually works.
The Results
When they tested RoDyGS on this new obstacle course and on real-world iPhone videos, it performed better than previous methods.
- The Outcome: It creates clearer, sharper 3D movies. If you watch a video of a spinning fan, RoDyGS can reconstruct the fan blades so they look solid and real, whereas older methods might make them look like a blurry smear or a ghost.
In summary: RoDyGS is a smart system that separates moving things from stationary ones, uses strict rules to keep moving objects from turning into digital ghosts, and figures out the camera's path on its own. This allows it to turn a simple, shaky phone video into a high-quality, 3D, time-traveling experience.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.