Forge4D: Feed-Forward 4D Human Reconstruction and Interpolation from Uncalibrated Sparse-view Videos
Forge4D is a feed-forward model that enables efficient, instant reconstruction and interpolation of dynamic 3D humans from uncalibrated sparse-view videos by jointly optimizing streaming 3D Gaussian reconstruction with self-supervised dense motion prediction and occlusion-aware fusion.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine trying to capture a person moving through a room and then instantly recreating that movement from any angle, even one no camera ever saw, or at a moment in time that was never recorded. This is the dream of four-dimensional reconstruction: building a living, breathing digital twin that exists not just in space, but in time. For years, achieving this has required either a studio full of synchronized cameras or hours of slow, computer-heavy processing that makes real-time interaction impossible. The goal has always been to turn a few scattered video clips into a complete, fluid 3D world where a person can be viewed from anywhere, at any moment, without the digital model falling apart or looking like a blurry mess.
A team of researchers has now taken a significant step toward making this dream a reality with a new system called Forge4D. Instead of trying to solve the entire problem of moving 3D shapes all at once, which often leads to confusion and errors, they broke the task down into two simpler, connected jobs. First, the system builds a static 3D model of the person and their surroundings from a few uncalibrated video streams. Think of this as creating a high-quality, frozen sculpture of the scene. Second, the system predicts exactly how every tiny part of that sculpture moves from one second to the next. By separating the shape from the motion, the researchers found a way to make the process fast enough to be useful in real time, even when the cameras recording the scene are not perfectly aligned or calibrated.
The core of their discovery lies in how they handle the data. Previous methods often tried to process an entire video sequence at once, which overwhelmed computer memory and slowed everything down. Forge4D takes a different approach by treating the video as a stream. It processes the frames one after another, using a special memory tool to remember what happened in the previous moments without needing to reload the entire history. This allows the system to maintain a consistent sense of scale and position as the person moves, ensuring that the digital model doesn't shrink, grow, or drift away as the video plays. It is a method that keeps the reconstruction stable and efficient, even when the input data is sparse and imperfect.
Once the system has a stable 3D model, it faces the challenge of predicting motion. Since there is no perfect map of how every point on a person's body moves in the real world to use as a teacher, the researchers invented a new way to teach the computer. They asked the system to guess the motion, then tried to warp the 3D model forward in time based on that guess. If the warped model looked like the actual next frame of the video, the guess was good. If it looked wrong, the system adjusted its understanding of the motion. This self-correcting loop, combined with a check against how light and shadow usually behave, allowed the computer to learn the complex dance of human movement without needing a pre-written script. It also developed a way to handle parts of the body that are hidden from view, ensuring that the digital model remains solid and does not flicker or glitch when a person turns or an object blocks the view.
The results of this approach are striking. When tested on various datasets, including complex scenes with people interacting with objects, the system produced images that were sharper and more realistic than those generated by previous methods. It could create new views of a scene from angles that were never filmed, and it could generate smooth, natural-looking movement for moments in time that were never captured. In tests, the system was able to produce these high-quality images in less than a quarter of a second per frame, a speed that opens the door to interactive applications like virtual reality, live broadcasting, and immersive communication. The researchers noted that while the system works exceptionally well for normal movements, it can struggle if the motion is extremely fast or if the time gap between frames is too large, suggesting that the assumption of smooth, continuous movement has its limits.
Ultimately, this work demonstrates that it is possible to reconstruct dynamic human scenes from simple, uncalibrated videos with a speed and quality that was previously out of reach. By simplifying the problem into manageable steps and teaching the computer to learn from its own predictions, the researchers have created a tool that brings us closer to a future where digital avatars can be generated instantly and interacted with in real time. The system does not just capture a moment; it understands the flow of time within that moment, allowing us to step into the video and look around as if we were truly there.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.