← Latest papers
💻 computer science

MeSS: City Mesh-Guided Outdoor Scene Generation with Cross-View Consistent Diffusion

This paper introduces MeSS, a novel framework that leverages city mesh models as geometric priors and employs a multi-stage image diffusion pipeline with cross-view consistency modules to generate high-quality, style-consistent outdoor scenes for virtual navigation and autonomous driving.

Original authors: Xuyang Chen, Zhijun Zhai, Kaixuan Zhou, Zengmao Wang, Jianan He, Dong Wang, Yanfeng Zhang, mingwei Sun, Rüdiger Westermann, Konrad Schindler, Liqiu Meng

Published 2026-08-05
📖 4 min read☕ Coffee break read

Original authors: Xuyang Chen, Zhijun Zhai, Kaixuan Zhou, Zengmao Wang, Jianan He, Dong Wang, Yanfeng Zhang, mingwei Sun, Rüdiger Westermann, Konrad Schindler, Liqiu Meng

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to build a perfect, walkable digital twin of a real city. You have the blueprints—the 3D shapes of every building, road, and tree—but they are all painted a flat, boring gray. They look like clay models waiting for a paint job. This is the world of "mesh models," which are becoming common for cities, but they lack the realistic textures (like brick, glass, and graffiti) that make a city feel alive. Without these textures, you can't use them for virtual reality tours or to train self-driving cars, because those systems need to see the world exactly as it looks, not just as a collection of shapes.

To fix this, scientists have tried using "generative AI," specifically a type of model called a "diffusion model." Think of these models as incredibly talented artists who can paint a picture from a simple description. However, when you ask these artists to paint a whole city street as you walk down it, they often get confused. They might paint a door on the left side of a building in one frame, and then suddenly paint it on the right side in the next frame, or change the color of the sky as you move forward. This is called "drift." The other big problem is that these artists often ignore the blueprints entirely, painting buildings that float in the air or don't match the actual shape of the city. The goal, then, is to teach an AI artist to paint a city that looks real, stays consistent as you walk through it, and perfectly hugs the 3D shapes of the buildings.

Enter MeSS (Mesh-Guided Scene Synthesis), a new method developed by researchers from Huawei, the Technical University of Munich, and others. Think of MeSS as a super-smart construction crew that doesn't just paint a picture; it builds a 3D world that you can actually walk through. Instead of guessing where the buildings are, MeSS starts with the city's "skeleton"—the textureless mesh model—and uses it as a strict guide.

The secret sauce of MeSS is a concept the authors call a "control-distribution match." Imagine you are teaching a robot to paint. If you train the robot using photos of real cities, it learns to paint based on how real cameras see the world. But if you then ask it to paint based on a computer-generated 3D model, it gets confused because the "depth" and "shadows" look different. MeSS solves this by training its "ControlNet" (the robot's instruction manual) directly on the computer-generated depth and shapes of the city mesh. This means the robot is an expert at reading the specific language of the 3D blueprints. When it starts painting, it doesn't guess; it knows exactly where the walls and windows should be because it was trained on the exact same kind of data it is using to paint.

To keep the city looking consistent as you travel down a long street (avoiding that "drift" problem), MeSS uses a clever two-step process. First, it acts like a master painter creating a series of "key frames"—important snapshots of the street every 20 meters. It paints these one by one, using the previous painting as a reference to ensure the style doesn't change. Then, it fills in the gaps between these snapshots with a technique called "Appearance Guided Inpainting." Imagine you have a blurry photo of a street corner; this tool looks at the sharp, clear parts of the photo around the blur and uses them to fix the blurry spots, ensuring the textures (like the pattern on a brick wall) match perfectly.

Finally, the team adds a "Global Consistency Alignment" step. Sometimes, even with good planning, one part of the street might look slightly brighter or darker than the next, like a photo that was edited in pieces. This step acts like a colorist in a movie studio, smoothing out those differences so the entire journey feels like one continuous, seamless video.

The results are impressive. On a test using a massive 200-meter city path, MeSS reduced visual inconsistencies by 31% and improved the overall quality of the generated images significantly compared to the best existing methods. Perhaps most excitingly, the researchers tested MeSS on a real-world city model of Munich (a "LoD3" mesh) without retraining the AI at all. The system successfully painted realistic facades that perfectly aligned with the real buildings' doors and windows, proving that this "mesh-guided" approach works even on real cities, not just computer simulations.

In short, MeSS shows that if you give an AI artist the right blueprints and teach it how to read them, it can build a virtual city that is not only beautiful but also geometrically perfect and consistent, opening the door for better virtual tours and safer self-driving car training.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →