Arbor: Explicit Geometric Conditioning for Controllable 3D Asset Generation
Arbor is a trainable attachment for text-conditioned latent 3D generation that introduces explicit geometric control via constraint meshes (hull, avoidance, and touch regions), enabling objects to satisfy specific spatial requirements while preserving visual quality and variation.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are trying to bake a cake using a magical recipe book (an AI). You tell the book, "Make me a chocolate cake," and it usually does a great job. But what if you need the cake to fit perfectly inside a specific box, or you need a hole in the middle for a candle, or you need the frosting to stop exactly where the plate begins?
Currently, if you try to explain these spatial rules with words ("make it fit the box"), the AI often gets confused. It might make the cake too big, ignore the hole, or spill frosting everywhere. You end up having to bake dozens of cakes, hoping one works, and then manually carving the bad ones to fix them.
Arbor is a new tool that solves this problem. It gives artists a way to draw the "rules of the room" before the cake is even baked.
Here is how it works, using simple analogies:
1. The "Traffic Light" Mesh
Instead of just typing a prompt, the artist creates a simple 3D shape (a mesh) that acts like a set of traffic lights for the AI. This shape has three specific colors (or "types") of regions:
- Green Zones (Hull): "The cake must exist here." (e.g., The seat of a chair).
- Red Zones (Avoidance): "The cake must not exist here." (e.g., The empty space above the seat so you can sit).
- Yellow Zones (Touch): "The cake must gently touch this surface, but not go past it." (e.g., The bottom of the chair legs touching the floor).
Think of this as drawing a blueprint where you don't draw the final furniture, but you draw the "ghosts" of where the furniture needs to be and where it needs to stay away.
2. The "Translator" (The Encoder)
The AI that makes the 3D objects speaks a secret, compressed language (latent space). The "traffic light" blueprint the artist draws is too big and detailed for the AI to read directly.
Arbor uses a special translator (a frozen geometric encoder) to shrink that big blueprint down into tiny, efficient "tokens" (like digital sticky notes). These notes carry the instructions: "Green here, Red there, Yellow here."
3. The "Smart Butler" (The Router)
This is the clever part. Imagine the AI is building the object piece by piece, like a construction crew working on different floors of a skyscraper.
- If the crew is working on the top floor, they don't need to know about the rules for the basement.
- If they are working on the legs, they don't need to know about the seat.
Arbor has a Smart Butler (a geometry router). It looks at which part of the object the AI is currently building and hands them only the relevant "sticky notes."
- If building the legs, the Butler hands over the "Touch the floor" note.
- If building the seat, the Butler hands over the "Must exist here" note.
This ensures the AI gets the right instructions for the right spot without getting overwhelmed by the whole blueprint at once.
4. The Result: A Reliable Baker
The paper tested this by asking the AI to make chairs, trucks, and sofas with specific rules.
- Without Arbor: The AI made nice-looking objects, but they often floated in mid-air, had holes where they shouldn't, or were too big for the space.
- With Arbor: The AI followed the rules perfectly. The chairs had legs that touched the ground, seats that fit the "Green Zone," and empty space where the "Red Zone" said to stay away.
Crucially, the AI didn't just copy the blueprint. It still used its creativity to make the chair look like a unique, stylish chair, not a boring copy of the drawing. It balanced obeying the rules with being creative.
Why This Matters
The authors call this "Explicit Geometric Conditioning." In plain English, it means giving the AI a clear, visual map of the physical space it needs to respect, rather than hoping it understands vague words.
The paper claims this makes 3D generation much more reliable for professional use (like video games), where objects need to fit specific spaces or interact correctly with other objects (like a character sitting on a chair without falling through it). It turns the "slot machine" approach of guessing prompts into a "co-authoring" process where the artist guides the AI with precise, spatial instructions.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.