AI vs Human Expert Reasoning: Assessing Agreements in Building Typology Predictions based on Street View Imagery
This study demonstrates that Vision-Language Models, particularly when enhanced with Chain-of-Thought prompting, can achieve approximately 70% accuracy in predicting building typologies from street view imagery by leveraging visual pattern recognition, though they differ from human experts by relying less on broader contextual cues and domain knowledge.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are trying to understand a city, but instead of walking its streets, you are looking at it through a giant, high-tech camera that sees everything from the ground up. This is the world of urban analysis, a field where scientists and planners try to map out how cities work, from the materials buildings are made of to how many floors they have. For a long time, to get this information, experts had to visit every single building, take notes, and guess what it was used for—a slow, expensive, and exhausting job. But recently, a new kind of "brain" has entered the game: the Vision-Language Model (VLM). Think of a VLM as a super-smart robot that has read almost every book on the internet and also learned to "see" pictures. It can look at a photo of a house and tell you not just what it looks like, but also what it is made of, how many stories it has, and what people do inside it, all by combining what it sees with what it knows about language. The big question everyone is asking is: Can this robot do the job as well as a human expert, like an architect or an engineer, especially in places where data is hard to find?
This paper takes that question to the bustling, chaotic, and vibrant streets of North Jakarta, Indonesia, a place where buildings are often built by hand, added to over time, and don't always follow official rules. The researchers wanted to see if these AI robots could look at street-view photos and correctly guess the "typology" (the fancy word for the type and use) of thousands of buildings. They set up a friendly competition between three of the most famous AI models—GPT-4o, Claude 3.5 Sonnet, and Gemini 2.0 Flash—and a team of human experts. The humans, who are civil engineers and architects, acted as the "gold standard," providing the correct answers based on their years of training. The AI models were given a special set of instructions, or "prompts," to see if asking them to "think step-by-step" (a technique called Chain-of-Thought) would help them get better grades.
The results were a mix of impressive highs and some funny, human-like mistakes. The AI models turned out to be pretty good at guessing the basics. When it came to counting how many floors a building had, the robots and the humans agreed about 70% to 80% of the time. It's like the AI is really good at counting the windows to guess the height. However, when it came to guessing what the building was used for (like a shop, a home, or a factory), the agreement dropped to around 60-70%. The AI sometimes got confused because a house and a small shop can look very similar from the street. The study found that the "Chain-of-Thought" method, where the AI is forced to explain its reasoning before giving an answer, made the models more stable and reliable, much like how a student does better on a test when they show their work.
One of the most interesting discoveries was how the AI and the humans thought differently. The AI was like a detective who only trusts what it can see with its eyes: it focused heavily on physical things like "walls," "windows," "materials," and "height." The human experts, on the other hand, were like seasoned locals who knew the neighborhood. They looked at the same pictures but also considered the "vibe" of the area, the condition of the building, and the context of the neighborhood. For example, if a building looked a bit worn out but was in a busy market area, a human might guess it's a shop, while the AI might just see "old bricks" and guess it's a warehouse. The paper suggests that while the AI isn't perfect yet, it can do about 70% of the work that humans do, and it can do it at a massive scale, looking at 30,000 buildings in a fraction of the time.
The researchers conclude that these AI tools are not ready to replace human experts entirely, especially for tricky jobs like guessing building use in messy, informal neighborhoods. However, they are powerful partners. They can act as a "first pass," quickly sorting through thousands of images to give planners a good idea of what a city looks like, which is a huge help in places where data is scarce or outdated. The study shows that by combining the AI's ability to spot visual patterns with human expertise in understanding context, we can build a much clearer picture of our cities. It's not about the robot winning; it's about the robot and the human working together to solve the puzzle of the urban world.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.