← Latest papers
💻 computer science

Foundation Model-Enabled Efficient Data Sampling (FEEDS): A label-efficient training strategy for pan-cancer, multi-tracer PET/CT datasets

The paper introduces FEEDS, a label- and compute-efficient training strategy that leverages foundation model embeddings to select the most informative unlabeled PET/CT cases for annotation, achieving fully-labeled-level performance across diverse cancer types and tracers with 70% less annotation effort.

Original authors: Biratal Raj Wagle, Bashirul Azam Biswas, Grant Chau, Matthew E. Maeder, Muhammad Azeem Arshad, Michael S. Leapman, James B. Yu, Indrani Bhattacharya

Published 2026-08-12
📖 4 min read☕ Coffee break read

Original authors: Biratal Raj Wagle, Bashirul Azam Biswas, Grant Chau, Matthew E. Maeder, Muhammad Azeem Arshad, Michael S. Leapman, James B. Yu, Indrani Bhattacharya

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to teach a robot to spot hidden treasures in a massive, chaotic attic. This isn't just any attic; it's a medical one, filled with thousands of 3D scans of human bodies. The "treasures" are cancer lesions, which can look like tiny specks or huge blobs, hiding in different organs and showing up differently depending on the type of scan used. To teach the robot, you need to show it examples of these treasures. But here's the catch: a human expert has to draw a perfect outline around every single treasure in every single scan. This is like asking a master artist to paint a detailed map of every grain of sand on a beach. It takes forever, costs a fortune, and there just aren't enough expert artists to go around.

For a long time, scientists tried to solve this by either showing the robot every picture (which is impossible because of the time it takes to draw the maps) or by letting the robot guess on the pictures it hasn't seen yet and hoping it learns from its mistakes (which often leads to the robot getting confused and making up fake treasures). The big question in this corner of science is: How do we teach a robot to be a master treasure hunter without making the human experts draw millions of maps? We need a way to pick the best few pictures to show the human, so that the robot learns the most from the least amount of work.

This is where a new strategy called FEEDS (Foundation Model-Enabled Efficient Data Sampling) comes in. Think of FEEDS as a super-smart librarian who has read every book in the world but hasn't seen the specific "treasure maps" yet. Instead of asking the librarian to read every single book to find the best ones to show the robot, FEEDS uses a special "vibe check" tool (called a foundation model) to look at the unmarked scans. This tool can instantly tell how different one scan is from another, like recognizing that one picture is a "sunny beach" and another is a "stormy forest."

The paper explains that FEEDS works in three simple steps. First, it takes all the unmarked scans and gives them a "vibe score" based on their visual features. Second, it looks at the small pile of scans the human experts have already marked (the "fixed labeled data") and asks, "Which unmarked scans are the most different from what we already know?" It specifically hunts for the weird, rare, or unusual cases that the robot hasn't seen before. Third, it asks the human experts to draw the maps only for those specific, unique cases. These new maps are then added to the robot's training pile, and the robot is trained just once.

The researchers tested this idea using a massive collection of cancer scans from different hospitals. They found that by using FEEDS to pick the most diverse and interesting cases, they could train a robot that was almost as good as one trained on every single scan in the database, but they only had to ask the human experts to draw maps for 30% of the data. In fact, the robot trained with FEEDS was better at finding real cancer spots and making fewer mistakes (like thinking healthy tissue was cancer) than robots trained by just picking random pictures. Even when they tested the robot on brand-new data from different hospitals it had never seen before, it still performed incredibly well, proving that it learned the "rules" of the game rather than just memorizing the specific pictures it was shown.

The authors suggest that this method is a game-changer because it saves a huge amount of time and effort. Instead of a slow, expensive process of labeling everything, or a confusing process of letting the robot guess and correct itself over and over, FEEDS offers a one-step, efficient path. It suggests that by being smart about which data we label, we can build powerful medical tools without overworking the doctors. The paper shows that this approach works across different types of cancer and different scanning machines, making it a practical tool for the future of medical imaging.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →