Frozen Brain-MRI Foundation Models Are Site Fingerprints
This paper reveals that frozen brain-MRI foundation models inherently encode acquisition site information as a dominant, linearly decodable fingerprint that often surpasses clinical variables, posing significant attribution risks and creating a trade-off between removing site bias and preserving anatomical segmentation accuracy.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are a detective trying to solve a mystery using a super-smart AI assistant. In the world of medical science, this AI is a "Foundation Model"—a massive, pre-trained brain that has already read millions of MRI scans of human heads. Scientists love these models because they are like having a genius medical student who has seen every type of brain imaginable, ready to help diagnose diseases or study anatomy without needing to be retrained from scratch. The big hope is that when you ask this AI to look at a new scan, it sees only the biology: the shape of the brain, the size of the ventricles, and the health of the tissue. It's supposed to be a pure window into the human body, ignoring everything else.
But here is the catch: medical scans aren't taken in a vacuum. They are taken in hospitals, using specific machines (scanners) that might be made by different companies, set to different speeds, or tuned to different settings. Just like how a photo taken with an iPhone looks different from one taken with a Samsung, even if it's the same person, MRI scans carry the "fingerprint" of the machine that took them. For years, scientists have worried that these machine differences might trick the AI, making it think a healthy brain from Hospital A is sick just because the picture looks slightly different from Hospital B. The big question was: Does the AI actually learn to see the brain, or is it secretly memorizing the style of the camera?
This paper investigates exactly that. The researchers took several of these "frozen" AI models—meaning they locked the AI's brain so it couldn't learn anything new—and tested what it was actually seeing. They treated the AI like a suspect and asked: "If I show you a brain scan, can you tell me which hospital it came from?" The answer was shocking. The AI didn't just see the brain; it saw the hospital with crystal-clear vision. In fact, the AI could identify the hospital (the "site") with about 90% accuracy, which was far better than it could identify the patient's age, sex, or even whether they had autism.
The most surprising twist? This wasn't because the AI had been "trained" to be a spy. The researchers tested this by using a version of the AI that had never been trained on any data at all—it was just a random collection of numbers. Even this "dumb," random AI could identify the hospital with nearly the same 90% accuracy. This means the "fingerprint" isn't a secret code the AI learned; it's baked into the raw pictures themselves. The way the machines capture the light and texture of the brain is so distinct that any system looking at the image, smart or dumb, can tell where it came from.
The paper also explored whether we could fix this. They tried to "scrub" the hospital information out of the AI's brain using mathematical tricks. They found that for simple tasks, like just looking at the whole brain and saying "yes" or "no," they could remove the hospital fingerprint without hurting the results. However, for detailed tasks like drawing a map of every tiny part of the brain, removing the fingerprint was disastrous. It turned out that the "hospital style" and the "brain anatomy" were tangled together in the AI's brain. When they cut out the hospital part, they accidentally cut out the tiny, deep structures of the brain too, like the thalamus and the hippocampus, making the AI unable to see them at all.
So, the main takeaway is a warning to scientists: Don't assume these powerful AI tools are purely seeing anatomy. They are also seeing the scanner. And because this "fingerprint" is so deeply rooted in the image itself, you can't just train the AI better to ignore it; you have to be very careful about how you use the AI's answers, especially when you are trying to map the tiny, intricate details of the human brain.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.