← Latest papers
💻 computer science

Looks Can be Deceiving: Annotator and Reviewer Performance Across Imagery Sources in Crowd-Sourced Aerial Damage Assessment

This paper empirically demonstrates that annotator and reviewer performance varies significantly across multi-source aerial imagery, with lower-resolution satellite data exhibiting substantially higher revision rates than higher-resolution drone or crewed aviation imagery, thereby challenging the efficacy of uniform quality-control workflows and suggesting the need for adaptive, source-aware curation strategies.

Original authors: Thomas Manzini, Priyankari Perali, Raisa Karnik, Stephen Johnson, Robin R. Murphy

Published 2026-08-18
📖 5 min read🧠 Deep dive

Original authors: Thomas Manzini, Priyankari Perali, Raisa Karnik, Stephen Johnson, Robin R. Murphy

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

When a disaster strikes, from a hurricane to a wildfire, the first priority is often to understand the scale of the damage. For decades, scientists and emergency responders have relied on cameras mounted on satellites, airplanes, and drones to take pictures of the affected areas from above. These images are powerful tools, but they are just raw data until humans look at them and decide what they are seeing. Is that building destroyed, or just damaged? Is that road blocked? To train the artificial intelligence systems that will one day help answer these questions automatically, researchers need vast libraries of these images, each one carefully labeled by a person. This process of teaching machines by showing them examples is the foundation of modern computer vision. However, a critical question has remained unanswered: does the source of the picture change how well a person can label it? A photo taken from a drone flying low offers a sharp, detailed view, while a satellite image taken from hundreds of miles away is much blurrier. It seems obvious that a clearer picture is easier to interpret, but until now, no one had measured exactly how much harder the blurry pictures make the job, or how much more likely a human is to make a mistake when the view is poor.

A team of researchers set out to solve this puzzle by examining a massive collection of images gathered after nine different disasters. They looked at a dataset containing nearly 75,000 buildings, each captured in three different ways: from a drone, from a crewed airplane, and from a satellite. The same buildings were visible in all three types of photos, allowing the team to compare how people performed on the exact same structures under different viewing conditions. To do this, they enlisted 187 volunteers, mostly high school and middle school students, to act as the human labelers. These students looked at the images and marked the damage level of each building using a standard scale. The work was not done in a single step; instead, it followed a rigorous process where the initial labels were checked by a single reviewer, and then a final committee of experts reviewed the entire set to reach a consensus on the correct answer. This multi-stage process gave the researchers a unique opportunity to see not just the final result, but how often the labels had to be changed at each step of the way.

The results revealed a striking pattern that challenges how these datasets are usually built. The researchers found that the clarity of the image had a direct and dramatic impact on the accuracy of the human labels. When the team compared the initial labels against the final, agreed-upon truth, they discovered that the error rate climbed steeply as the image quality dropped. For the high-resolution drone images, the final committee had to change about 6.85 percent of the labels. For the medium-resolution images from crewed airplanes, that number jumped to 14.05 percent. But for the low-resolution satellite images, the committee had to revise nearly 37 percent of the labels. This means that for every three buildings labeled from a satellite photo, one of them was likely misidentified by the initial worker. The same trend held true at every stage of the review process; the lower the resolution, the more the labels needed to be corrected.

Perhaps even more surprising was the finding that a single person checking the work was not enough to fix these errors, especially for the blurry images. In a typical workflow, a second person might review the first person's work to catch mistakes. The researchers found that this single review did help, but it left a significant amount of disagreement unresolved. After the first reviewer had looked at the satellite images, the final committee still had to change more than 20 percent of the labels. For the drone images, the committee still had to change nearly 7 percent. This suggests that when the image is unclear, a single second opinion is often insufficient to reach the right answer. The data showed that the reviewers were not necessarily making different mistakes than the original workers; rather, they were often failing to intervene at all when the image was too confusing to be sure. It was only when a group of people came together to discuss and deliberate on the difficult cases that the labels finally aligned with the consensus.

These findings suggest that the current methods for creating these large image libraries may be inefficient and risky. If researchers treat all images the same, giving them the same amount of review time and attention, they are likely leaving the most error-prone data—the blurry satellite photos—unfixed. This could lead to artificial intelligence models that are biased or unreliable when they encounter real-world situations where only low-quality images are available. The researchers argue that to build better systems, the review process must be tailored to the source of the image. Instead of a uniform approach, more effort should be directed toward the lower-resolution sources, and the final decision on difficult cases should rely on a group consensus rather than a single reviewer. By understanding that the human eye struggles more with certain types of views, scientists can design better workflows that ensure the data feeding our future AI tools is as accurate as possible, regardless of where the camera was flying.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →