When Experts Disagree: Characterizing Annotator Variability for Vessel Segmentation in DSA Images
This paper analyzes and quantifies the variability among multiple expert annotators segmenting cranial blood vessels in 2D DSA images to characterize segmentation uncertainty, with the goal of guiding future annotations and developing uncertainty-aware automatic segmentation methods.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are trying to draw a map of a tiny, twisting city inside a human head, but you can only see it through a foggy window. This is the world of medical imaging, specifically a technique called Digital Subtraction Angiography (DSA). Doctors use DSA to take pictures of blood vessels in the brain, which look like a complex web of roads. To help computers learn how to read these maps automatically, humans have to draw the lines of the roads first. This is called "segmentation." Usually, we assume that if a human expert draws the line, that drawing is the perfect "truth." But what if two experts look at the same foggy window and draw slightly different lines? What if one sees a tiny alleyway that the other misses? This paper dives into that exact mystery. It asks a simple but tricky question: If even the experts can't agree on where the blood vessels are, how can we teach a computer to get it right? The answer isn't just about better drawing; it's about understanding that the "truth" might be a little bit fuzzy, and that's okay.
The researchers behind this study decided to stop pretending that there is only one perfect way to draw these brain vessels. They gathered 66 real images of brain blood vessels from patients and asked two experts to trace the vessels by hand. One expert was a doctor-in-training specializing in blood vessel procedures, and the other was a scientist who knows how to teach computers to see. As expected, the two experts didn't draw the exact same lines. Sometimes they agreed perfectly, but often they disagreed, especially on the tiniest, thinnest vessels that are less than 1 or 2 pixels wide. To figure out who was "right," the team didn't just pick a winner. Instead, they brought in a whole panel of 11 other judges—doctors and scientists—to look at 100 small pieces of the images and vote on whether a vessel was there or not.
The results were fascinating and a bit surprising. When both original experts agreed a vessel was there, the new judges agreed with them 90% of the time. But when the experts disagreed, the judges were split. If Expert A said "vessel" and Expert B said "no vessel," the judges sided with Expert A 76% of the time. However, if the roles were reversed, they only sided with Expert A 16% of the time. This showed that the "ground truth" isn't a single, solid line; it's more like a cloud of possibilities. The team also measured how much the experts' drawings overlapped using a special math tool called "clDice," which checks if the center of the drawn lines matches up. They found that while the experts were mostly in sync, there was still a noticeable gap, especially for the thinnest vessels.
To fix the problem of measuring these tiny, fuzzy lines, the researchers invented a new way to calculate the scores. The old way was too strict; if an expert's line was just a tiny bit off from the other's, the score would crash, even if they were both looking at the same thin vessel. The team created a "Modified clDice" that acts like a safety net. Instead of demanding the lines match pixel-for-pixel, it asks, "Is the line close enough?" They tested this with different "distance thresholds," like a fuzzy border of 2.5 pixels around the line. This new method showed that the experts were actually in much better agreement than the old, strict math suggested.
So, what does this all mean? The paper doesn't claim to have solved the problem of perfect vessel drawing. Instead, it suggests that we need to change how we teach computers. We shouldn't just feed them one person's drawing and say, "This is the truth." Instead, we should teach them to understand the uncertainty. By using the disagreement between experts as a clue, we can tell the computer, "Hey, this area is tricky; maybe we need more eyes on it," or we can build AI models that know when to be careful. The team also shared their work with the world: a database of about 2,000 image patches with all these different drawings, and the open-source software they used to measure the uncertainty. In short, they proved that in the messy world of medical imaging, admitting that experts disagree is actually the first step toward building smarter, more reliable tools.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.