Feed-Forward Gaussian Splatting from Sparse Aerial Views
This paper introduces AnyCity, a feed-forward generative framework that reconstructs large-scale urban scenes from sparse aerial views by combining observation-grounded geometry with scaffold-conditioned diffusion priors to resolve evidence imbalances and eliminate artifacts like ghosting and melted facades.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are trying to build a 3D model of a whole city, but you only have a few blurry photos taken from a drone flying high up.
The Problem: The "Roof vs. Wall" Blind Spot
When a drone flies straight down or at a shallow angle, it sees the tops of buildings (roofs) and open streets very clearly. It sees them over and over again. But the sides of the buildings (facades), the back of distant skyscrapers, and anything hidden behind other structures are almost invisible.
If you try to build a 3D city from just these few photos, standard computer programs get confused. They might turn a solid brick wall into a melted puddle, stretch a building's texture like taffy, or leave gaps where the walls should be. This happens because the computer is trying to guess the missing parts based on very little evidence.
The Solution: AnyCity
The researchers created a new tool called AnyCity. Think of it as a smart construction crew that knows exactly which parts of the city are real and which parts need a little creative help.
Here is how it works, step-by-step:
1. The "Solid Scaffold" (The Observation-Supported Part)
First, AnyCity looks at the photos and builds a "scaffold" or a skeleton using only the parts it can clearly see.
- Analogy: Imagine a carpenter building a house frame. They nail down the beams and walls that are clearly visible in the blueprints. They don't guess where the windows go yet; they just secure the parts they are 100% sure of.
- Result: This creates a stable, reliable base for the roofs and open areas, ensuring the computer doesn't "hallucinate" fake structures where it has good data.
2. The "Creative Fill-in" (The Completion Tokens)
Once the solid scaffold is built, AnyCity turns its attention to the missing parts—the invisible walls and distant buildings.
- Analogy: Now, the carpenter brings in a "ghostwriter" or a creative architect. This expert looks at the solid frame and says, "Okay, I know there's a wall here because that's how cities usually look, even though I can't see it in the photo."
- How it works: The system uses a "video prior" (a brain trained on thousands of hours of city videos) to guess what the missing textures and shapes should look like. But crucially, it doesn't just replace the whole building with a guess. It only adds a "residual update"—a gentle correction to the weak spots.
3. The "Safety Gate" (Observation Preservation)
This is the most important trick. The system has a strict rule: "Do not change what you can already see."
- Analogy: Imagine the creative architect is trying to fix the invisible wall, but they are wearing a helmet that locks their hands if they try to touch the solid beams the carpenter already built. They can only work on the empty spaces.
- Result: This prevents the "ghostwriter" from accidentally melting the real walls or changing the shape of the roof just because it thinks it knows better. It keeps the real data real and only fills in the blanks.
4. The "Teacher-Student" Training
To teach this system, the researchers used a clever training method.
- The Teacher: A computer that gets to see many photos of the city (dense views). It knows the city perfectly.
- The Student: The AnyCity model, which only sees the few photos (sparse views).
- The Lesson: The Teacher shows the Student what the city looks like when fully observed. The Student tries to guess the missing parts to match the Teacher's view, but it must keep the parts it can already see exactly the same. This teaches the Student how to fill in the blanks without messing up the known facts.
The Final Result
When you give AnyCity just a few drone photos, it spits out a complete, coherent 3D city model in seconds.
- The Good: The roofs and roads look exactly like the photos. The invisible walls and distant buildings look realistic and consistent, not melted or stretched.
- The Speed: Unlike older methods that take hours to calculate, AnyCity does this in a single "feed-forward" pass (like reading a sentence from start to finish without stopping to think).
In Summary:
AnyCity is like a master builder who builds a house based on a few photos. It nails down the parts it can see with absolute precision, then uses its knowledge of how cities usually look to gently fill in the invisible parts, all while making sure it never accidentally changes the parts it already built. This results in a realistic 3D city that is both accurate to the photos and complete in its structure.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.