A Comparative Evaluation of Structural MRI Foundation Models for Age, Sex, and Body-Mass Index Predictions
This study presents the first systematic benchmark of four structural MRI foundation models, demonstrating that while performance varies across tasks and datasets, models like 3D-Neuro-SimCLR and AnatCL generally outperform traditional FreeSurfer-based baselines in predicting age, sex, and body-mass index, thereby establishing a new standard for evaluating learned representations in neuroimaging via the BrainFMBench platform.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer
The human brain is a complex organ, and understanding how it changes over time or differs between people is a central goal of modern neuroscience. To see inside the living brain without surgery, researchers rely on magnetic resonance imaging, a technique that creates detailed three-dimensional maps of brain structure. For decades, scientists have analyzed these maps by manually measuring specific features, such as the thickness of the brain's outer layer or the size of its folded surface. These measurements, known as morphometric features, have provided a solid foundation for studying conditions like Alzheimer's disease or the natural effects of aging. However, collecting enough brain scans to study rare conditions or diverse populations is difficult, and traditional methods often require a great deal of human effort to interpret the data. In recent years, a new approach has emerged: training computer systems on massive collections of brain scans so they can learn to recognize patterns on their own. These systems, called foundation models, are designed to understand the raw images directly, potentially offering a more powerful way to extract meaningful information from medical scans than older methods.
A team of researchers set out to test whether these new computer systems truly offer an advantage over the established methods used for years. They focused on four publicly available models that had been pre-trained on large datasets of brain images. To see how well these models worked in real-world scenarios, the researchers gathered magnetic resonance images from three different groups of people: those with Parkinson's disease, a network of healthy children and adults, and a large cohort of healthy volunteers. They asked the computer models to perform three specific tasks that are common in brain research: determining a person's sex, estimating their age, and predicting their body mass index. To ensure a fair comparison, the team tested these new models against a standard approach that relies on measuring the thickness and surface area of the brain's cortex, a method that has long been the gold standard in the field. They used a consistent testing framework where the computer models were not allowed to learn anything new during the test, ensuring that the results reflected only what the models had already learned from their initial training.
The results showed that the new computer models did not automatically win every time. While some of the foundation models performed better than the traditional measurements on specific tasks or with certain groups of people, the overall picture was mixed. Two of the models, 3D-Neuro-SimCLR and AnatCL, consistently outperformed the traditional methods across most of the tests. Among them, 3D-Neuro-SimCLR showed the most reliable performance, though it did struggle with one specific task involving sex classification in the HBN dataset. The other models tested did not consistently beat the traditional measurements, suggesting that the ability of a computer to learn its own patterns from raw images does not guarantee better results in every situation. The researchers also found that the way these new models understood the brain was different from how traditional measurements described it, indicating that the computer was seeing relationships in the data that human-designed measurements might miss.
This study provides a clear benchmark for the field, showing that while these advanced computer models hold great promise, they are not yet a universal replacement for established techniques. The findings suggest that models like 3D-Neuro-SimCLR and AnatCL are particularly strong tools for improving predictions in brain imaging, but their success depends heavily on the specific question being asked and the population being studied. To help other scientists continue this work, the researchers made their testing methods available to the public, creating a living resource where new models can be added and compared as they are developed. This transparency ensures that the field can move forward with a clear understanding of which tools work best, avoiding the assumption that newer technology is always superior without proof. The work confirms that these foundation models are a valuable addition to the neuroscientist's toolkit, offering a path to more accurate predictions when the right model is matched to the right task.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.