LaS-Comp: Zero-shot 3D Completion with Latent-Spatial Consistency
LaS-Comp is a zero-shot, training-free framework that leverages 3D foundation model priors through a two-stage design of explicit replacement and implicit refinement to achieve high-quality, category-agnostic 3D shape completion, validated by the new Omni-Comp benchmark.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are an artist trying to finish a sculpture, but you only have a few broken shards of clay in your hand. Your goal is to recreate the entire statue, making sure the new parts fit perfectly with the old shards and look like a real, complete object.
This is the challenge of 3D Shape Completion. For a long time, computers struggled with this. If you showed them a broken chair, they might guess it's a table, or they might smooth over the broken parts so much that the chair looks like a blob. They often needed to be trained on thousands of "broken chair + whole chair" pairs, which limits them to only what they've seen before.
Enter LaS-Comp, a new "super-artist" AI that can finish any 3D object it sees, even if it has never seen that specific object before. Here is how it works, explained simply:
The Problem: The "Translation" Glitch
The paper starts with a clever observation. Modern AI models that generate 3D shapes work in a "secret code" (called Latent Space). Think of this like a language where the AI speaks.
- The Issue: If you take a broken chair and ask the AI to translate it into its secret code, and then take a whole chair and translate that, the "secret code" for the matching parts (the legs, the seat) doesn't actually match! It's like if you wrote "cat" in English and "cat" in French, but the AI thought they were completely different words.
- The Result: If the AI tries to just "fill in the blanks" using this secret code, it gets confused and creates a mess because the code for the broken part doesn't line up with the code for the whole part.
The Solution: LaS-Comp's Two-Step Dance
The authors created a framework called LaS-Comp (Latent–Spatial Consistency). Instead of trying to fix the secret code directly, they use a two-step dance to keep the original shape safe while letting the AI do its magic.
Step 1: The "Sticky Tape" (Explicit Replacement Stage)
Imagine you have a broken vase. Before asking the AI to guess the rest, you take the actual broken shards and physically tape them back onto the new vase you are about to build.
- What the AI does: It takes the "secret code" of the new shape and forcibly swaps out the parts that correspond to your broken shards with the actual data from your shards.
- Why it helps: This guarantees that the AI cannot accidentally change your original input. The broken chair leg stays exactly where it is. It's like saying, "I know this part is real; don't touch it."
Step 2: The "Smoothie Blender" (Implicit Alignment Stage)
Now, you have a vase where the taped-on shards look a bit jagged against the new clay. There's a rough seam.
- What the AI does: It performs a quick, tiny "polish." It looks at the edge where the real shard meets the new clay and gently nudges the math to make the transition smooth. It's like using a fine sandpaper to blend the new clay into the old shard so you can't tell where one ends and the other begins.
- Why it helps: This prevents the "glitchy" look where the new parts look like they are floating or disconnected from the old parts.
Why This is a Big Deal
- Zero-Shot (No Training Required): You don't need to teach this AI with thousands of examples. It uses a "pre-trained" brain (a 3D foundation model) that already knows what chairs, cars, and animals generally look like. It just applies that knowledge to your specific broken piece.
- Works on Anything: Whether you have a single photo of a chair, a random chunk of a robot, or a missing piece of a statue, this method handles it. Previous methods often failed if the broken piece was too weird or if the object was partially hidden.
- Speed: It's incredibly fast, finishing a shape in about 20 seconds.
The New "Test Drive" (Omni-Comp)
The authors realized that old tests were too easy. They mostly used simple "single photo" scans. To prove their method is truly tough, they built a new test called Omni-Comp.
- Imagine a driving test. Old tests were like driving on a straight, empty road.
- Omni-Comp is like driving in a storm, with potholes, random obstacles, and missing road signs. It tests the AI on:
- Random Crops: Like someone took a bite out of the object.
- Semantic Parts: Like showing only the wheels of a car, with the rest missing.
- Real-World Noise: Messy, real-life scans from robots and cameras.
The Verdict
In simple terms, LaS-Comp is like a master restorer who doesn't just guess what the missing piece looks like. Instead, they:
- Lock the original broken pieces in place so they don't move.
- Grow the rest of the object around them using their deep knowledge of how the world works.
- Polish the seams until the whole thing looks like it was made in one piece.
The result? A computer that can look at a broken toy, a half-scanned car, or a missing statue part, and instantly "hallucinate" the rest of it with high accuracy, without needing to be taught how to do it first.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.