← Latest papers
🤖 AI

Resolution Meets Reduction: Efficient Visual Context for 3D Radiology Report Generation

This paper systematically investigates the optimal allocation of visual token budgets across input resolution, field of view, and compression strategies for 3D radiology report generation, demonstrating that anatomy-guided region-of-interest cropping and high-resolution inputs paired with PerceiverResampler projectors significantly improve clinical performance on CT datasets.

Original authors: Jonathan Suprijadi, Raphael Stock, Moritz Langenberg, David Zimmerer, Kim-Celine Kahl, Stefan Denner, Yannick Kirchhoff, Karol Gotkowski, Maximilian Rokuss, Jeremias Traub, Tassilo Wald, Constantin Ul
Published 2026-08-11
📖 4 min read☕ Coffee break read

Original authors: Jonathan Suprijadi, Raphael Stock, Moritz Langenberg, David Zimmerer, Kim-Celine Kahl, Stefan Denner, Yannick Kirchhoff, Karol Gotkowski, Maximilian Rokuss, Jeremias Traub, Tassilo Wald, Constantin Ulrich, Klaus Maier-Hein

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to teach a brilliant but very hungry robot to read medical scans and write doctor's reports. This robot is a "Vision-Language Model," a type of AI that can look at pictures and understand words. The challenge is that a 3D CT scan (a detailed, three-dimensional X-ray of the inside of a body) is like a library of millions of tiny image tiles. If you show the robot every single tile, it gets overwhelmed, like a student trying to read an entire encyclopedia in one second. To fix this, scientists usually try to either shrink the picture (lowering the resolution) or cut out the parts they don't need (cropping). But there's a tricky trade-off: if you shrink the picture too much, the robot misses tiny but important clues, like a small fracture or a tiny spot. If you keep the picture huge, the robot runs out of "brain power" and can't finish the report. The big question is: how do you feed the robot just the right amount of information so it stays smart but doesn't get a headache?

This paper, titled "Resolution Meets Reduction," is like a massive cooking competition where the chefs are trying to find the perfect recipe for feeding these AI robots. The researchers tested different ways to prepare the "ingredients" (the CT scans) and different "chefs" (the AI models) to see which combination produces the best medical report. They didn't just guess; they cooked up thousands of variations using two huge datasets of real CT scans and reports. They found that the secret isn't just about making the picture smaller or bigger, but about how you compress the information.

The most consistent trick they discovered is "anatomy-guided cropping." Think of it like this: instead of showing the robot a photo of a whole house to find a broken window, you zoom in specifically on the window. By cutting out the empty sky and the lawn and focusing only on the lungs or the abdomen, the robot can see the tiny details much more clearly without needing more brain power. This simple step improved the robot's accuracy in almost every test they ran.

However, simply zooming in isn't always enough. The researchers also tested different "compressors"—tools that squish the huge amount of image data down into a smaller package the robot can digest. They found that some compressors are like a bad photocopier that blurs everything when you shrink it, while others are like a smart editor that keeps the important details sharp even when the file size is tiny. Specifically, they found that two types of compressors, called "TokenPacker" and "PerceiverResampler," were the best at keeping the fine details alive even when the data was squeezed down by 64 times.

The paper also tested different "brains" (the Large Language Models) to see if a bigger brain meant a better report. Surprisingly, the size of the brain mattered less than the quality of the eyes (the Vision Encoder) and the quality of the compression. A slightly smaller brain with a great camera and a smart compressor often beat a giant brain with a blurry camera.

In the end, the team built the best possible setup they could find. On one test dataset, their best robot achieved a score of 49.5, and on another, it hit 49.0. These scores are the highest ever recorded for this specific task, beating previous record-holders. The study suggests that if we want AI to help doctors write reports, we shouldn't just throw more computing power at the problem. Instead, we should be smarter about how we crop the images and how we compress the data, ensuring the robot sees the tiny, critical details that could save a life.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →