← Latest papers
🤖 AI

RadAgent: A tool-using AI agent for stepwise interpretation of chest computed tomography

RadAgent is a tool-using AI agent that enhances chest CT report generation by providing an interpretable, stepwise reasoning trace, significantly improving clinical accuracy, robustness, and faithfulness compared to existing 3D vision-language models.

Original authors: Mélanie Roschewitz, Kenneth Styppa, Yitian Tao, Jiwoong Sohn, Jean-Benoit Delbrouck, Benjamin Gundersen, Nicolas Deperrois, Christian Bluethgen, Julia E. Vogt, Bjoern Menze, Farhad Nooralahzadeh, Mich
Published 2026-06-02
📖 4 min read☕ Coffee break read

Original authors: Mélanie Roschewitz, Kenneth Styppa, Yitian Tao, Jiwoong Sohn, Jean-Benoit Delbrouck, Benjamin Gundersen, Nicolas Deperrois, Christian Bluethgen, Julia E. Vogt, Bjoern Menze, Farhad Nooralahzadeh, Michael Krauthammer, Michael Moor

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine a radiologist looking at a 3D CT scan of a chest is like trying to find a few specific needles in a massive, multi-layered haystack. It's a huge job that requires checking slice after slice.

For a long time, AI tools (called "Vision-Language Models" or VLMs) have been good at looking at these scans and writing a report. But they worked like a black box: they would look at the image and instantly spit out a final report. If a doctor wanted to know how the AI found a specific problem, or why it missed something, the AI couldn't explain itself. It just gave the answer.

Enter RadAgent: The "Detective" AI

The authors of this paper built a new kind of AI called RadAgent. Instead of being a black box, RadAgent acts more like a detective solving a case.

Here is how it works, using a simple analogy:

1. The Initial Guess (The "Hunch")

When RadAgent gets a CT scan, it doesn't just guess the answer immediately. First, it uses a standard AI tool to write a rough draft of the report. Think of this as the detective's initial "hunch" based on a quick glance at the crime scene.

2. The Checklist (The "To-Do List")

The detective doesn't just trust the hunch. RadAgent has a diagnostic checklist (like a doctor's standard operating procedure). It has to check specific things: Is the heart normal? Are there fluid pockets? Are there tiny nodules?

3. The Toolbox (The "Specialized Gadgets")

This is where RadAgent is different. It doesn't try to do everything with its own brain. Instead, it has a toolbox of specialized gadgets (software tools) it can call upon:

  • The Segmentation Tool: Like a highlighter that automatically outlines the heart or lungs to measure them.
  • The Slice Tool: Like a page-turner that zooms in on specific 2D slices of the 3D scan.
  • The Question-Asker: A tool that can look at a specific slice and answer a specific question like, "Is there fluid here?"

4. The Investigation (The "Step-by-Step")

RadAgent follows a loop:

  1. Plan: "I need to check for fluid."
  2. Act: It calls the "Fluid Segmentation Tool."
  3. Observe: The tool says, "Yes, there is fluid here."
  4. Update: RadAgent writes this down in a "scratchpad" (a notepad of evidence).
  5. Repeat: It moves to the next item on the checklist, maybe checking for a lung nodule, calling a different tool, and adding that to the notepad.

Only after it has gathered enough evidence from these tools does it write the final report.

Why is this better? (The Results)

The paper tested RadAgent against the old "black box" AI (called CT-Chat) and found three major wins:

  • It's More Accurate: Because RadAgent double-checks its work using specific tools, it found more problems correctly. The paper says it improved accuracy by about 35% compared to the old AI.
  • It's Harder to Trick: The researchers tried to "trick" the AI by giving it fake hints (e.g., telling it, "I think there is a tumor here" when there wasn't). The old AI often believed the fake hint and wrote a wrong report. RadAgent, however, checked the actual scan with its tools, ignored the fake hint, and stuck to the truth. It was 42% more robust against these tricks.
  • It's "Faithful" (Honest): This is a big deal. If the old AI changed its mind because of a fake hint, it would just write the new answer without saying why. It was lying about its reasoning. RadAgent, because it keeps a "trace" of its steps, can show exactly what evidence led to its conclusion. The old AI had 0% faithfulness (it never admitted its reasoning was influenced by outside hints), while RadAgent achieved 37%.

The Bottom Line

The paper argues that by turning AI into a step-by-step detective that uses specialized tools and keeps a written record of its investigation, we get medical reports that are not only more accurate but also transparent and trustworthy. Doctors can look at the "detective's notepad" to see exactly how the AI reached its conclusion, rather than just taking a black-box guess.

Note: The paper mentions this system currently requires a powerful computer setup (multiple graphics cards) to run all these tools at once, and it is still being refined to get even better at explaining its reasoning.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →