Text-Image Conditioned 3D Generation
This paper introduces TIGON, a dual-branch framework for text-image conditioned 3D generation that leverages the complementary strengths of visual fidelity from images and semantic guidance from text to overcome the limitations of single-modality approaches.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you want to build a custom 3D statue for your living room. You have two ways to tell the builder what you want:
- The Photo: You hand them a picture of a specific chair.
- The Description: You write a note saying, "I want a red velvet chair with gold legs."
For a long time, 3D AI builders could only listen to one of these instructions at a time. And that caused problems.
The Problem: The "One-Track" Builder
- If you only gave the Photo: The builder would copy the chair perfectly from the angle you showed them. But if you showed them the front of the chair, they would have no idea what the back looked like. They would just guess (or "hallucinate"), often making the back look weird or wrong.
- If you only gave the Description: The builder would understand that you want a "red velvet chair," but they wouldn't know the specific style. They might build a red chair, but it could look like a cheap plastic toy instead of the fancy velvet one you imagined.
The result? Either the shape was right but the details were wrong, or the details were right but the shape was a mess.
The Solution: The "Dual-Brain" Builder (TIGON)
The paper introduces a new system called TIGON. Think of TIGON as a construction team with two specialized brains working together:
- The Visual Brain: This brain is obsessed with the photo. It says, "Okay, I see the texture, the color, and the exact shape of the front. I'll make sure the 3D model looks exactly like that."
- The Language Brain: This brain is obsessed with your text. It says, "I see you said 'gold legs' and 'red velvet.' I'll make sure the back of the chair matches that description, even though we can't see it in the photo."
How they work together:
Instead of fighting, these two brains talk to each other constantly while building.
- If the photo shows a weird angle, the Language Brain steps in to say, "Don't guess! The text says it's a tiger tail, so make it a tail, not a random blob."
- If the text is vague, the Visual Brain says, "The text just says 'chair,' but the photo shows a specific vintage style. Let's copy that style."
They use a clever trick called "Zero-Initialization." Imagine the two brains are connected by a door that is locked at the start. At first, they work separately. As they learn, the door slowly unlocks, allowing them to share just the right amount of information without getting confused.
Why This Matters
The researchers tested this by showing the AI a "low-quality" photo (like a blurry side view of a toaster) and asking it to build a 3D toaster.
- Old AI: Built a toaster that looked nothing like the one in the picture because it couldn't guess the missing parts.
- TIGON: Built a perfect toaster. It used the blurry photo to get the general shape and the text description ("silver toaster with two slots") to fill in the missing details perfectly.
The Big Takeaway
This paper proves that combining sight and language is the secret sauce for high-quality 3D generation. It's like having a builder who can both see exactly what you want and understand exactly what you're saying, resulting in 3D objects that are not only beautiful but also exactly what you asked for.
In short: TIGON is the first 3D artist that doesn't have to guess when you give it a partial photo; it just asks your text description for the rest of the story.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.