← Latest papers
💻 computer science

Can Large Language Models Substitute for Human Raters? Reliability and Validity of Urban Design Quality Assessment

This study evaluates the reliability and validity of multimodal LLMs for urban design assessment in Seoul, finding that while they perform well on visually explicit features, their inability to fully replicate human ratings for contextual or proportional variables limits their suitability to element-specific rather than universal applications.

Original authors: Wookjae Yang, Jihyun Hwang, Reid Ewing

Published 2026-06-26
📖 4 min read☕ Coffee break read

Original authors: Wookjae Yang, Jihyun Hwang, Reid Ewing

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to grade the beauty and safety of a neighborhood. Traditionally, you'd hire three expert inspectors to walk the streets, look at the buildings, count the trees, and fill out a detailed report card. This is accurate, but it takes a long time and costs a lot of money.

Now, imagine you have three super-smart AI robots (called Multimodal Large Language Models, or LLMs) that can look at the same street photos and fill out the exact same report cards instantly. The big question this paper asks is: Can we just fire the human inspectors and let the robots do the whole job?

The short answer from the study is: Not yet. The robots are great at some things, but they miss the mark on others.

Here is a breakdown of what the researchers found, using simple analogies:

1. The "Big, Obvious" vs. The "Small, Subtle"

Think of the neighborhood like a messy room.

  • What the Robots Get Right: If you ask the robots to count the big, obvious items—like "How many red brick walls are there?" or "Are there any large commercial signs?"—they are surprisingly consistent. They are like a robot that can perfectly count the number of red chairs in a room. They see the big, structural things clearly.
  • What the Robots Miss: If you ask them to estimate "How much of the wall is covered by windows?" or "How much of the sky can you see above the trees?", they get confused. They also struggle to spot tiny, scattered items like a single trash can on the sidewalk or a security camera hidden in a corner. It's like asking the robot to estimate the exact percentage of a rug that is blue versus red, or to find a single lost coin in a pile of leaves. They often miss these small details or guess the proportions wrong.

2. The "Agreement" Problem

The researchers had three human inspectors and three AI robots look at the same 40 spots in a Seoul neighborhood.

  • Humans vs. Humans: The three humans mostly agreed with each other. If one said a street was "safe," the others usually agreed.
  • Robots vs. Robots: The three robots also mostly agreed with each other. They were consistent in their own way.
  • Humans vs. Robots: This is where it gets tricky. Even though the robots were consistent with themselves, they often disagreed with the humans.
    • The Analogy: Imagine the humans and robots are both taking a photo of a sunset. The humans might say, "It's 70% orange." The robots might say, "It's 30% orange." They are both looking at the same thing, but they are measuring it differently. The robots tend to see fewer streetlights, fewer trees, and fewer signs than the humans do.

3. The "Comfort" Test

To see if the robots' reports actually mattered, the researchers checked if their scores matched how comfortable people felt on those streets.

  • The Good News: For some things, like "Are there trees?" or "Is there a fence?", the robots and humans agreed on the pattern. When the robots said "more trees," people felt more comfortable, just like when the humans said "more trees."
  • The Bad News: For other things, the robots got it wrong. For example, the humans thought "parking under a building" made people feel more comfortable, but the robots didn't see that connection. The robots also missed the link between "noise" and "discomfort" that the humans caught.

The Bottom Line

The study concludes that you can't just swap humans for AI in urban design yet.

  • The Robots are "Specialists," not "Generalists." They are excellent at spotting big, clear architectural features (like building materials or signs).
  • They are "Blind" to Context. They struggle with things that require guessing proportions, spotting small details, or understanding how different parts of a street work together to create a feeling.

The Verdict: Instead of replacing the human inspectors, the best approach right now is a hybrid team. Let the robots do the easy, obvious counting (like counting big signs), but keep the humans to handle the tricky stuff (like estimating tree coverage or spotting small safety hazards). The robots are a helpful assistant, but they aren't ready to take the lead.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →