← Latest papers
📄 medicine

How errors from prompt-assisted longitudinal tumour segmentation change RECIST 1.1 target-lesion response categories: a retrospective out-of-fold analysis

This retrospective out-of-fold analysis of 258 patients demonstrates that prompt-assisted longitudinal tumor segmentation errors altered RECIST 1.1 response categories in approximately 17% of cases, with measurement errors being a significant but not uniquely dominant contributor compared to detection and localization failures.

Original authors: Souraj Adhikary, Negar Chabi, Andre Mastmeyer

Published 2026-08-10
📖 5 min read🧠 Deep dive

Original authors: Souraj Adhikary, Negar Chabi, Andre Mastmeyer

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are a detective trying to solve a mystery: is a villain getting weaker or stronger? In the world of cancer treatment, the "villain" is a tumor, and the "detectives" are doctors using special cameras called CT scanners to take pictures of the inside of a patient's body. To know if the medicine is working, doctors measure the size of specific tumors at the start of treatment and then again a few weeks later. They use a strict rulebook called RECIST 1.1 to decide the verdict. If the tumor shrinks enough, the patient gets a "Partial Response" (good news!). If it grows, it's "Progressive Disease" (bad news!). If it stays the same, it's "Stable Disease."

But here is the tricky part: measuring a squishy, irregular blob inside a 3D body is hard. Even a tiny mistake in drawing the outline of the tumor—like drawing a line just a few millimeters too wide or too narrow—can flip the verdict from "shrinking" to "growing." Recently, scientists have started using computer programs (Artificial Intelligence) to help draw these outlines automatically, hoping to make the job faster and more consistent. But the big question remains: if the computer makes a mistake, does it actually change the doctor's final verdict? Does a tiny error in the drawing trick the computer into thinking the patient is getting better when they aren't, or vice versa?

This paper investigates exactly that. The researchers took a group of 258 patients with metastatic melanoma (a type of skin cancer that has spread) and used a smart computer program called "LongiSeg" to try and redraw the tumor outlines on follow-up scans. The computer was given a "hint" (a click) showing where the tumor was, but it had to do the actual drawing itself. The researchers then compared the computer's drawing against the "gold standard" drawing made by human experts.

The results were a bit surprising. In about 16.7% of the patients (that's roughly 43 out of 258, or one in six), the computer's drawing led to a different verdict than the expert's drawing. For example, the expert might say the patient's condition is "Stable," but the computer's slightly different measurement would say "Partial Response" (getting better), or the other way around.

The researchers then played detective to figure out why these mistakes happened. They broke the errors down into four types:

  1. Omission: The computer completely missed the tumor (like forgetting to draw a character in a comic).
  2. Mislocalization: The computer drew the tumor in the wrong place (like drawing a character on the wrong side of the page).
  3. False Persistence: The tumor was actually gone, but the computer kept drawing it (like drawing a ghost that isn't there).
  4. Sizing Error: The computer found the right tumor but drew it slightly too big or too small.

The study found that Sizing Errors were the biggest culprit, contributing to 8.3% of the disagreements. However, the researchers were careful to note that this wasn't a clear, overwhelming victory. When they combined the "missed" and "wrong place" errors into one group, that group was almost as big as the sizing errors. In fact, the difference between them was so small that it could have been due to chance. This means that while getting the size right is crucial, simply finding the tumor in the first place is just as important.

Interestingly, the computer made mistakes that made the patient look better than they actually were in 28 cases, and made them look worse in 15 cases. But the authors warn us not to get too excited about this "good news" bias. When they changed the rules slightly (like counting a different number of tumors), this pattern shifted. So, they can't say for sure that the computer is secretly trying to make patients look better; it just seems to lean that way in this specific test, but it's not a proven rule.

The study also ran some computer simulations to see if these tiny drawing errors were enough to cause the verdict flips. They found that yes, even a small error of just 2 mm in the diameter of the tumor could flip the result. This confirms that the computer needs to be incredibly precise, not just in finding the tumor, but in measuring its edges perfectly.

In the end, the paper concludes that while AI is a helpful tool, it isn't perfect yet. About one in six patients might get a different "report card" if we rely solely on this specific computer program's measurements. The study highlights that we can't just focus on whether the AI finds the tumor; we also have to worry about how accurately it measures it. The authors stress that this is just one specific model tested on one type of cancer, and more testing is needed before we can trust these digital detectives to make the final call on a patient's health.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →