Using street view images and visual LLMs to predict heritage values for governance support: Risks, ethics, and policy implications
This paper evaluates the use of multimodal Large Language Models to analyze street view images for identifying heritage values in Sweden's building stock to support National Building Renovation Plans, while critically addressing the associated risks, ethical concerns, and policy implications regarding transparency, error detection, and sycophancy in governance.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are trying to renovate an entire city's houses to save energy, but you have a massive blind spot: you don't know which houses are special "heritage" buildings that need extra care. In Sweden, there is no complete list of these special homes. It's like trying to pack a suitcase for a trip without knowing which items are fragile antiques and which are just old t-shirts.
This paper describes a project where Swedish authorities tried to solve this problem using a new kind of "digital eye" called a Multimodal Large Language Model (LLM). Here is how they did it, what they found, and the warnings they discovered along the way.
The Digital Detective
Instead of sending human experts to walk down every street in Sweden (which would take forever and cost a fortune), the researchers used an AI trained to "see" and "read." They fed the AI over 150,000 photos of building facades taken from Google Street View.
Think of the AI as a very fast, very well-read art student. The researchers gave it a specific checklist (based on how human experts usually judge buildings) and asked: "Look at this photo. Does this building look like it has historical or cultural value? Rate it from 1 to 100."
They didn't teach the AI new rules; they just asked it to use its existing knowledge to make a "zero-shot" guess. It's like asking a knowledgeable friend to judge a painting without showing them a textbook first.
The Results: A Good Start, But Not Perfect
The AI managed to scan a huge area of the building stock (about 5 million square meters of floor space) and flag about 1-2% of the buildings as "potentially heritage."
- The Good News: The AI was surprisingly good at spotting the obvious "old and fancy" buildings. If a building had classic ornaments or looked very old, the AI gave it a high score.
- The Bad News: The AI was a bit "conservative." It tended to ignore modest, everyday buildings that might still be culturally important to a local community. It also seemed to have a bias: it loved buildings in wealthy areas and in Stockholm (the capital) more than buildings in poorer areas or smaller towns.
The "Sycophant" Problem
One of the most interesting findings in the paper is about the AI's personality. The researchers call this "sycophancy."
Imagine a yes-man who always agrees with you and speaks very confidently, even when he's wrong. LLMs are trained to be polite and well-spoken. If you ask them a question, they will give you a very confident, detailed answer, even if they are hallucinating (making things up).
The paper warns that if authorities just trust the AI's confident text, they might get fooled. To fix this, the researchers forced the AI to output numbers (like a score of 50 or 80) instead of long paragraphs of text. This made it easier for human experts to check the math and spot errors, rather than being impressed by the AI's smooth talking.
The Human Safety Net
The paper emphasizes that this AI tool is not a replacement for human experts. It's more like a high-speed sorting machine.
- The Process: The AI sorts through millions of photos and flags the ones that might be special.
- The Human Role: Human heritage experts then look at the flagged buildings to make the final decision.
The researchers tested this by having experts look at a random sample of the AI's top picks. The experts agreed that the AI had successfully found buildings that did have heritage features. However, they also noted that the AI missed many other buildings that humans would have considered valuable.
The Big Warnings
The paper ends with a few crucial "caution signs" for anyone using this technology:
- Bias is Real: The AI learned from data created by humans, and humans have biases. The AI tended to think buildings in rich neighborhoods were more "heritage-worthy" than those in poorer ones. If used blindly, this could lead to unfair policies.
- Context Matters: A building's value isn't just about how it looks from the street. It's about its history, who lives there, and what it means to the community. The AI can only see the "skin" of the building (the facade), not the "soul" (the history).
- Don't Trust the "Yes-Man": Because the AI is so confident, humans need to be extra careful not to let it make the final call. It is a tool to help, not a tool to decide.
The Bottom Line
This project showed that AI can be a powerful "assistant" for governments, helping them quickly scan a country to find buildings that might need protection. However, it is not a magic wand. It works best when used as a first step to narrow down a huge list, leaving the final, nuanced decisions to human experts who understand the local history and culture.
In short: Use the AI to find the needles in the haystack, but let the human expert decide which needles are actually precious.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.