← Latest papers
💬 NLP

Multi-Agent AI System for Radiology Report Structuring and Quality Assurance with Independent Radiologist Evaluation

This study demonstrates that a locally deployed multi-agent AI system effectively structures radiology reports into standardized anatomical sections and performs quality assurance checks with favorable performance as validated by independent radiologist evaluation.

Original authors: Iryna Hartsock, Cesar Lam, Christopher Otteni, Aliya Qayyum, Robert Gatenby, Cyrillo Araujo, Ghulam Rasool

Published 2026-08-19
📖 5 min read🧠 Deep dive

Original authors: Iryna Hartsock, Cesar Lam, Christopher Otteni, Aliya Qayyum, Robert Gatenby, Cyrillo Araujo, Ghulam Rasool

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

In the world of medical imaging, a radiologist's report is a critical bridge between a scan and a patient's care. When a doctor orders a CT scan of the chest, abdomen, or pelvis, a specialist must translate the images into a written narrative describing what they see. These reports are vital, yet they are often written in free-flowing text, where one doctor might list findings by body part while another groups them by the order they noticed them. This lack of a standard structure can make it difficult for other doctors to find specific details quickly, potentially leading to missed information or wasted time searching for answers. Furthermore, because these reports are dictated into speech-to-text systems, they can contain subtle errors, such as a finding in the lungs being described in the conclusion as if it were in the liver, or a critical warning that was never explicitly communicated to the treating team. As hospitals generate thousands of these reports daily, ensuring they are both consistent and error-free has become a significant challenge that traditional manual review struggles to meet.

To address this, a team of researchers at the Moffitt Cancer Center in Tampa, Florida, developed a new kind of artificial intelligence system designed to organize and check these reports without changing the words the doctors actually wrote. Instead of rewriting the reports or summarizing them into a brief list, this system acts like a meticulous librarian who takes a chaotic stack of handwritten notes and files every single sentence into the correct drawer, while simultaneously scanning the whole document for contradictions. The team built a "multi-agent" system, which means they created a small team of specialized computer programs that work together. One part of the system uses strict, pre-written rules to sort sentences into anatomical categories like "Lungs" or "Liver." When the rules are not enough to decide where a sentence belongs, a more advanced reasoning program steps in to understand the context and make the call. This process happens entirely on the hospital's own secure computers, keeping patient data private.

The researchers tested this system on 638 real radiology reports from CT scans performed in 2023 and 2024. These reports were created by 15 different board-certified radiologists, each with their own unique writing style. The system successfully took the unstructured "Findings" section of every report, which contained a total of 22,270 sentences, and rearranged them into a standardized format. For example, a sentence describing a heart issue that was buried in the middle of a paragraph about the lungs was moved to the "Heart" section, while a sentence about a liver lesion was placed under "Liver." Crucially, the system did not delete, summarize, or invent any text; it simply moved the original sentences into their proper places. The entire process took an average of about 56 seconds per report, with the most time-consuming part being the reasoning step for the more complex sentences.

Beyond just organizing the text, the system also acted as a quality control inspector. It scanned the reports for specific types of errors, such as when the "Findings" section described a problem that the "Impression" section at the end of the report failed to mention, or when a finding contradicted the patient's gender. In this group of 638 reports, the system flagged 90 of them, or about 14 percent, as containing potential inconsistencies. The most common issue it found was a mismatch between what was described in the body of the report and what was summarized at the end. It also caught rare instances where a finding was critical but not clearly communicated, or where a sentence seemed to describe an organ that did not match the patient's anatomy.

To verify if the system was actually helpful, two independent radiologists reviewed a sample of 45 reports that had been processed by the AI. They looked at whether the sentences were moved to the right sections and whether the error flags were accurate. The doctors agreed that in 69 percent of the cases, the AI had restructured the report correctly. In only 4 percent of the cases did they find a clear mistake where a sentence was put in the wrong place. In the remaining cases, the doctors disagreed with each other on where a sentence should go, suggesting that some medical findings can reasonably belong to more than one category. Importantly, both reviewers confirmed that the system never left out any important medical information and never made up facts that were not in the original report. When asked about the quality of the error checking, the doctors rated the system's performance as "excellent" or "good" in 84 percent of the reports they reviewed.

The study suggests that combining a rule-based sorter with a reasoning-based checker can create a reliable workflow for standardizing medical reports. While the system did not make the reports perfect, and the doctors noted that the restructured versions were sometimes just as good as the originals rather than significantly better, the system proved it could handle the varied writing styles of 15 different doctors without losing any content. The researchers found that the system was particularly good at catching mismatches between different parts of a report, a task that is difficult for humans to do consistently when reviewing hundreds of reports a day. Although the study was limited to a specific set of CT scans and a relatively small group of doctors for the final review, the results indicate that such a system could support radiologists by providing a consistent structure and a safety net for common reporting errors, all while preserving the original voice and details of the medical professional who dictated the report.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →