AnyStyle: Single-Pass Multimodal Stylization for 3D Gaussian Splatting
AnyStyle is a feed-forward framework that enables zero-shot, pose-free 3D Gaussian Splatting reconstruction and stylization by integrating multimodal (textual and visual) conditioning into existing backbones, offering superior style controllability while preserving geometric quality.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are an architect trying to build a 3D model of a real-world room, but you only have a few blurry photos taken from random angles, and you don't know exactly where the camera was standing when each photo was snapped. For a long time, building a perfect 3D world from such messy clues was like trying to solve a giant jigsaw puzzle while blindfolded; it took computers hours of grinding math to figure out the shape of the walls and the furniture. But recently, a new technique called "3D Gaussian Splatting" changed the game. Think of this method as painting a scene not with solid bricks, but with millions of tiny, fluffy, colored clouds (Gaussians) that float in space. When you look at them from different angles, they blend together to look like a solid, realistic object, and they do it incredibly fast. This is a big deal because it means we can finally create 3D worlds for video games or movies in seconds instead of hours.
However, there was a catch. While these new methods could build the shape of the room quickly, they couldn't easily change its look to match a specific artistic style, like turning a photo of a living room into a Van Gogh painting or a cyberpunk city. Previous attempts to do this were either too slow, required the computer to re-learn the whole room from scratch for every new style, or could only copy a style if you showed it a picture, leaving no room for you to say, "Make it look more like a watercolor painting with soft edges." This is where the new paper, "AnyStyle," steps in to fix the problem.
The researchers behind AnyStyle have built a clever new system that lets you take those messy, unposed photos of a scene and instantly turn them into a 3D world with a completely new artistic style, all in a single, lightning-fast pass. The magic trick is that you don't just have to show the computer a picture of the style you want; you can also just type a description, like "make it look like a sketch with heavy charcoal strokes" or "turn it into a neon-lit cyberpunk scene." The system understands both pictures and words, mixing them together to guide the 3D clouds into the right shapes and colors.
To understand how they pulled this off, imagine the computer's brain is split into two teams. The first team is a "frozen" expert that never changes; its only job is to look at your photos and figure out the geometry—where the walls, floors, and objects are. This team is reliable and fast because it was already trained on thousands of scenes. The second team is the "artist," which is free to experiment. This team takes the shape information from the first team and the style instructions (from your text or image) and paints the scene. The genius part of their design is a special "injection" mechanism. Instead of forcing the artist to relearn how to see the world, they simply slip the style instructions into the artist's workflow like a secret note. This note tells the artist exactly how to tweak the colors and textures without messing up the shape of the room. Because the shape expert is frozen and the artist is lightweight, the whole process happens in a fraction of a second—specifically, in under 0.1 seconds for each input image.
The paper shows that this method is a significant improvement over previous attempts. While other systems that try to do this in one go often struggle to follow complex instructions or require the computer to train from scratch every time, AnyStyle works with a pre-trained model and just adds a small, flexible layer on top. The researchers tested their system by asking people to compare the results with other top methods. The results were clear: people preferred the AnyStyle images because they looked more like the intended art style while keeping the original scene's details intact. The system also proved it could handle both text and images equally well, allowing users to refine a style by mixing a reference photo with a text prompt like "make the colors softer."
In short, the authors found that by separating the job of "figuring out the shape" from the job of "applying the style," and by using a smart way to inject style instructions that doesn't break the existing system, they can create high-quality, stylized 3D worlds instantly. They didn't just suggest this might work; they measured it, showing that their method produces better results than current state-of-the-art tools and does so without needing hours of training time. This means that in the near future, you might be able to take a few photos of your bedroom and instantly see what it would look like as a watercolor painting, a comic book scene, or a futuristic city, all without waiting for a computer to crunch the numbers for hours.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.