How Many Visual Levers Drive Urban Perception? Interventional Counterfactuals via Multiple Localised Edits
This paper proposes a lever-based interventional counterfactual framework that uses structured, localized image editing to identify which specific visual elements—such as mobility infrastructure or physical maintenance—most significantly influence human perceptions of urban safety.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
The "Urban Remote Control": Making Cities Feel Safer, One Click at a Time
Imagine you are walking through a neighborhood at night. You feel a sudden prickle of unease. Is it because the streetlights are dim? Is it because there’s graffiti on the walls? Or is it because the sidewalk is cracked and messy?
If you could grab a magic remote control for the city, which button would you press to make that street feel instantly safer? Would you add a flowerbed, repaint the crosswalk, or fix a broken storefront?
That is the question researchers at University College London (UCL) are trying to answer.
The Problem: The "Black Box" of City Feelings
Right now, we have AI models that can look at a photo of a street and say, "This place feels unsafe" or "This place feels wealthy." These models are good at spotting patterns, but they are terrible at explaining why.
It’s like having a friend who says, "I don't like this restaurant," but when you ask why, they just shrug. They can tell you what they feel, but they can't tell you exactly what needs to change to make them like it.
The Solution: The "Lever" Framework
The researchers created a way to turn that "shrug" into a precise list of instructions. Instead of just looking at pixels, they created "Visual Levers."
Think of these levers like the settings on a soundboard in a music studio. Each lever controls one specific "vibe" of the city:
- The Maintenance Lever: Fix the graffiti, clean the litter, or repair the walls.
- The Greenery Lever: Add some plants or trees.
- The Infrastructure Lever: Repaint the lines on the road or fix the crosswalks.
How the "Magic Remote" Works (The Process)
The researchers built a three-step digital factory to test these levers:
- The Architect (The Planner): The AI looks at a real photo and says, "Okay, in this specific street, I could try to 'pull the lever' to add greenery near that entrance or remove that trash pile."
- The Artist (The Generator): An AI artist takes the original photo and "edits" it based on those ideas. It’s like using Photoshop, but instead of a human doing it, the AI tries to "imagine" what the street would look like if the lever were pulled.
- The Inspector (The Auditor): This is the most important part. Because AI can sometimes get "hallucinations" (like accidentally turning a tree into a giant purple monster), a second AI acts as a strict inspector. It checks: "Did we actually just add greenery, or did we accidentally change the whole street? Does this look real, or does it look like a cartoon?" If the edit fails the inspection, it gets thrown in the trash.
What Did They Find? (The Results)
They tested this on 50 different scenes from cities like San Francisco, Singapore, and Amsterdam. Even though they used an AI "proxy" to judge safety rather than real humans (which is their next step), they found some fascinating patterns:
- The "Clean Up" Effect: Fixing physical things—like repairing a building's face or cleaning up litter—consistently made the AI think the street looked safer.
- The "Roadway" Boost: Surprisingly, the biggest "safety boost" came from Mobility Infrastructure. Simply repainting lane markings or crosswalks made the streets look much more "orderly" and safe.
- The "Greenery" Paradox: Adding plants made things look better, but trying to "manage" trees (like pruning them) actually made the AI think the street looked less safe. This might be because the AI thinks "more leaves = more hiding spots," even if a human would see it as "neatly trimmed bushes."
Why Does This Matter?
This isn't just about playing with photos. This research is a blueprint for the future of urban planning.
Instead of city officials guessing, "Maybe we should plant more trees?" they could use these "levers" to run simulations. They could ask: "If we spend \10,000 on repainting crosswalks versus \10,000 on cleaning graffiti, which one will actually make our citizens feel more secure walking home at night?"
It’s about moving from guessing how a city feels to engineering how a city feels.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.