SimuScene: Simulation-Ready Compositional 3D Scene Reconstruction from a Single Image
SimuScene introduces a novel compositional 3D reconstruction pipeline that integrates a physics engine directly into the generative process to diagnose and correct geometric errors like interpenetration and instability, thereby enabling the creation of stable, simulation-ready 3D scenes from a single image for robotic manipulation tasks.
Original authors:Inhee Lee, Sangwon Baik, Sungjoo Kim, Hyeonwoo Kim, Hyunsoo Cha, Hanbyul Joo
Original authors: Inhee Lee, Sangwon Baik, Sungjoo Kim, Hyeonwoo Kim, Hyunsoo Cha, Hanbyul Joo
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). ✨ This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
The Problem: The "Ghost" Furniture
Imagine you take a photo of a messy desk. You want to turn that photo into a 3D world where a robot can walk around, pick up a cup, or push a book.
Current technology is like a very talented but slightly clumsy artist. If you show them a photo, they can guess what the hidden parts of the objects look like (like the back of a bookshelf). However, they often make mistakes:
The Ghost Cup: They might draw a cup floating in mid-air because they couldn't see the table underneath it.
The Solid Wall: They might draw a chair that is partially inside the table, as if the table and chair are made of the same ghostly material.
The Collapse: If you put these "3D guesses" into a physics simulator (a digital sandbox that follows the laws of gravity), the scene falls apart. The floating cup drops, the chairs sink through the floor, and the whole digital room collapses into a pile of junk.
The Solution: SimuScene (The "Physics Detective")
The researchers at Seoul National University created SimuScene. Instead of just guessing what the 3D shapes look like and hoping for the best, they built a system that uses physics as a detective to fix its own mistakes while it is building the scene.
Think of it like this:
The First Draft (The Guess): The system looks at the photo and creates a rough 3D sketch of the objects, just like other methods do.
The "Drop Test" (The Detective): Before finalizing the scene, the system drops these objects into a digital gravity simulator.
If a cup falls through a table, the simulator says, "Hey! That table is too short or the cup is floating!"
If two chairs are stuck inside each other, the simulator says, "They are overlapping! Push them apart!"
The Correction (The Fix): This is the magic part. The system doesn't just ignore these errors. It uses the distance the objects fell or the amount they overlapped as a clue.
Stretching: If a chair leg is too short and the chair falls, the system stretches the leg until it touches the floor.
Resampling: If the shape is completely wrong (like a cup that looks like a cube), the system throws that shape away and asks the AI to "re-draw" the object, this time using the physics clues to guide the new drawing.
The "Physics-in-the-Loop" Analogy
Imagine you are trying to build a tower of blocks based on a blurry photo.
Old Way: You build the tower, step back, and realize the top block is floating. You then try to tape it to the air. It looks weird and unstable.
SimuScene Way: As you place each block, you gently tap the tower. If a block wobbles or falls, you immediately know that block is the wrong size or shape. You swap it out for a better one before you finish the tower. You use the wobble as a signal to fix the shape.
What Makes It Special?
The paper highlights three main tricks:
Sequential Building: It builds the scene one object at a time, starting from the bottom (the floor) and working up. This prevents objects from pushing each other into impossible positions.
Smart Stretching vs. Re-drawing: If an object is just a little off (like a slightly short leg), it stretches the mesh. If the object is totally wrong (like a hidden part of a sofa), it uses a special "resampling" technique to generate a new, better shape.
The "Wall" Check: It uses a smart AI (a Vision-Language Model) to ask, "Is this picture frame hanging on the wall, or is it standing on a shelf?" This ensures hanging objects stay attached to the wall and don't fall to the floor.
The Result
When the researchers tested this, their 3D scenes were stable.
When they dropped the objects in a simulator, they didn't sink or float.
They stayed exactly where the photo showed them.
They could be used immediately for robot tasks, such as teaching a robot arm how to pick up a bottle or teaching a humanoid robot how to walk through a cluttered room, without the robot crashing into invisible walls or falling through the floor.
In short, SimuScene turns a single photo into a 3D world that obeys the laws of physics, not just the laws of art. It uses gravity as a teacher to correct its own mistakes in real-time.
Technical Summary: SimuScene
Problem Statement
Reconstructing interactive, simulation-ready 3D scenes from a single image is a critical bottleneck for robotic manipulation and embodied AI. While recent "single-image lifters" (e.g., SAM3D) can recover plausible per-object shapes, composing them into a full scene often results in configurations that collapse under physical simulation. These failures manifest as interpenetrating meshes, objects hovering without support, or sinking into surfaces due to errors in occlusion-induced shape completion and monocular pose estimation. Existing physics-aware methods typically treat physics as a post-hoc layout correction tool, adjusting object positions after reconstruction but leaving underlying geometric errors unresolved. Consequently, when the geometry itself is incorrect, layout adjustments alone cannot produce a physically plausible configuration.
Methodology: SimuScene
SimuScene is a compositional 3D reconstruction pipeline that integrates physics directly into the loop of shape and layout estimation. Rather than using physics merely for cleanup, the method utilizes the physics engine as a diagnostic measurement tool during the generative process to drive geometric corrections.
The pipeline operates through the following stages:
Decomposed Scene Initialization:
The input image is processed to separate static base structures (e.g., tables, walls) from movable objects.
Base structures are reconstructed first to serve as fixed colliders.
Remaining objects are lifted using SAM3D to generate initial meshes, poses, and scales.
Semantic tags are assigned to objects (free, point-anchored, or line-anchored) using Vision-Language Models (VLMs) to define their physical constraints (Degrees of Freedom).
Pose Refinement (Pre-Simulation):
Initial poses from SAM3D are refined against image evidence (depth and segmentation) using FoundationPose to minimize rotational and translational errors before simulation, preventing severe penetration artifacts.
Diagnostic Simulation (Physics as a Probe):
Objects are processed sequentially in ascending order of their position along the gravity axis.
Penetration Resolution: Inter-object overlaps are resolved by displacing the active object along a simulator-informed direction.
Gravity-Based Diagnostics: Objects are released under gravity. The resulting displacement (Δt) and rotation (ΔR) serve as diagnostic signals.
Metric: A normalized gravity-axis displacement (ρi) is calculated. Small deviations (ρi<0.15) indicate minor accumulated errors, while large deviations suggest fundamental shape-sampling failures due to severe occlusion.
Physics-Informed Shape Correction:
Gravity-Axis Stretch: For minor errors, the canonical mesh is stretched along the gravity-aligned axis to correct support failures without resampling.
Amodal Shape Resampling: For severe errors, the method triggers a resampling process. It constructs an auxiliary "amodal" crop mask by augmenting the visible mask with the projected extent of the desired bounding box. This crop is fed back into a fine-tuned SAM3D to generate a more complete 3D shape.
Fine-Tuning: SAM3D is fine-tuned using a Direct Preference Optimization (DPO) objective with flow-matching. The model is trained on synthetic pairs contrasting occlusion-failed latents with completed amodal latents, using LoRA adapters to preserve the base model's capabilities while improving occlusion robustness.
Final Composition:
Corrected objects are assembled, and a final sequential penetration resolution is performed.
A full settling simulation is run where all free objects settle jointly under gravity to produce a stable, simulation-ready scene.
Key Contributions
Physics-in-the-Loop Diagnostic Simulation: The authors introduce a sequential, per-object protocol that integrates physical dynamics directly into reconstruction. Physical violations (e.g., interpenetration, gravity-induced displacement) are converted into actionable diagnostic signals rather than being treated as post-hoc artifacts.
Physics-Informed Shape Correction: A two-tier geometry update mechanism addresses violations via gravity-axis stretching for minor errors and OBB-guided amodal resampling for severe shape failures. This directly resolves geometric errors that layout-only methods miss.
Extensive Experimentation: The paper provides comprehensive evaluations across diverse datasets (GraspClutter6D, Aria Digital Twin, GenWild) and demonstrates that the reconstructed scenes support downstream robotics applications, including humanoid control and robot-arm manipulation.
Experimental Results
SimuScene achieves state-of-the-art performance on physical stability and geometric alignment benchmarks:
Physical Stability: The method significantly reduces penetration ratios and mean displacement errors compared to baselines like SAM3D, Gen3DSR, and 3D-RE-GEN. For instance, on the GenWild dataset, SimuScene achieved a penetration ratio of 2.5% compared to 16.2% for SAM3D.
Geometric Alignment: The method maintains high alignment with the input image view even after physics simulation, avoiding the "collapse" or "floating" artifacts common in other methods.
Amodal Resampling Quality: In pairwise evaluations using VLMs, the resampled meshes generated by the fine-tuned model achieved a substantially higher Elo score and win rate compared to raw SAM3D outputs and simple stretching operations, confirming superior recovery of occluded geometries.
Downstream Utility: The reconstructed scenes were successfully deployed in:
Humanoid Control: Training physics-based character policies for dexterous human-object interaction (HOI).
Robot-Arm Manipulation: Serving as input for closed-loop VLM agents to execute text-guided 6D pose prediction and grasping in cluttered environments.
Significance and Claims
The paper claims that SimuScene reframes physical simulation from a post-hoc validator into an in-the-loop diagnostic signal. By inverting physical violations into geometric corrections, the method resolves the structural ambiguities inherent in monocular perception. The authors assert that this approach produces scenes that are directly usable in downstream tasks such as robotic manipulation and reinforcement learning without manual scene authoring. While acknowledging that the sequential protocol cannot revise early estimates based on later evidence, the paper posits that integrating physical dynamics as a generative supervisory signal offers a scalable foundation for reconstructing simulation-ready scenes from single images, with joint scene-level optimization identified as a natural next step.