CFB-GBM v2.0: An Augmented Longitudinal Dataset for Multi-Modal Glioblastoma Segmentation, Radiomics, and RANO Progression Tracking
CFB-GBM v2.0 is an enhanced, publicly available longitudinal dataset of 264 glioblastoma patients that features nearly complete (97%) multi-modal Gross Tumour Volume delineations across all timepoints, derived RANO 2.0 progression labels, and pre-computed radiomic features to facilitate reproducible research in treatment response prediction and personalized medicine.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Glioblastoma is the most aggressive form of primary brain cancer in adults, a disease that moves with frightening speed and leaves patients with a median survival of only fifteen months. Because the tumor behaves so differently from one person to the next, doctors struggle to predict who will respond to treatment and who will not. Currently, the standard way to judge if a therapy is working relies on magnetic resonance imaging, or MRI, taken over several months. Doctors look for changes in the size of the tumor, but this process is slow. It often takes two months to tell the difference between a tumor that is actually growing and one that only looks larger because of inflammation caused by the treatment itself. This delay is critical; in a race against time, waiting two months to adjust a failing treatment plan can mean the difference between life and death. To solve this, researchers need vast amounts of detailed medical data that track patients over time, showing exactly how their tumors change and how they react to therapy. However, such data has been scarce, often locked away in private hospital records or missing crucial details needed to build reliable computer models.
A team of researchers from France has now released a significant update to a public database designed to help scientists and doctors overcome these hurdles. Called CFB-GBM v2.0, this new collection brings together the medical records and scans of 264 patients who were treated for glioblastoma at a single oncology center between 2017 and 2023. The dataset is special because it follows these patients through three distinct moments in their treatment journey: before therapy begins, and then again at roughly four and six months afterward. For each of these moments, the database includes a rich variety of images, such as different types of MRI scans and CT scans, along with precise maps of the radiation doses the patients received. The most important addition in this new version is that the researchers have now filled in the missing pieces of the puzzle: they have created detailed outlines of the tumors for every single patient at every single time point. In the original version of the data, these outlines were missing for most patients at the later stages, making it impossible to study how the tumors changed over time. Now, with the outlines complete for nearly all patients, the database allows researchers to see the full story of the disease's progression.
To create these new outlines, the team did not manually trace every single tumor by hand, a task that would have taken years. Instead, they trained a sophisticated computer program to do the work. They started with a model that had already learned to recognize brain tumors from a different, large collection of medical images. They then taught this program using the few high-quality outlines they already had from their own patients. Once the computer learned the patterns, it automatically generated the missing tumor outlines for the rest of the group. To ensure the computer was not making mistakes, five expert radiation oncologists reviewed a random selection of the new outlines. They found that the computer's work was remarkably accurate, matching the experts' own corrections almost perfectly. This validation gives other scientists confidence that they can use these automatically generated outlines to train their own tools without needing to spend months redrawing the tumors themselves.
With the tumor outlines now complete for every patient at every stage, the researchers were able to calculate exactly how much the tumors shrank or grew between visits. They applied a standard set of rules used by doctors worldwide to classify the results. These rules categorize a patient's response as a complete disappearance of the tumor, a partial shrinkage, a stable condition where the tumor does not grow significantly, or a progression where the tumor grows larger. By doing this for every possible pair of time points, the team has provided a clear, ready-to-use record of how each patient responded to treatment. This makes the dataset immediately useful for researchers trying to build artificial intelligence systems that can predict treatment success early on, potentially allowing doctors to switch strategies before the disease advances too far.
Beyond the tumor outlines and response categories, the dataset includes other tools designed to make research easier and more consistent. The team has provided pre-calculated measurements of the tumor's texture and shape, known as radiomic features. These numbers describe the internal structure of the tumor in ways the human eye cannot see, offering clues about the biology of the cancer. They have also included masks that isolate the brain from the rest of the head in the images, which helps researchers focus their analysis on the relevant area without interference from bone or skin. Furthermore, the database now clearly notes which version of the World Health Organization's classification guidelines was used to diagnose each patient, a detail that matters because the definition of glioblastoma changed in 2021 to require specific genetic markers.
Despite these major improvements, the researchers are careful to note what the data cannot yet tell us. The database does not include information about the specific genetic mutations inside the tumors, such as the status of the IDH gene, which is a powerful predictor of how a patient will fare. Without this genetic information, the dataset cannot be used to study the link between a tumor's genes and its behavior. The authors suggest that future updates could include these missing genetic details and break the tumor outlines down into even finer parts, such as separating the dead center of the tumor from the active edges. For now, however, CFB-GBM v2.0 stands as a robust and accessible resource. It transforms a collection of scattered medical records into a coherent, longitudinal story of 264 patients, offering the scientific community a solid foundation to develop better ways to track, predict, and ultimately treat this devastating disease.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.