Self-supervision drives representational convergence in medical foundation models more than clinical supervision
This study reveals that representational convergence in medical foundation models is primarily driven by self-supervised pretraining objectives rather than clinical supervision or model scale, resulting in modest alignment that supports cross-encoder transfer learning but fails to fully capture clinical judgment or semantic similarity.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are trying to teach a group of very different robots to understand the world. You give them all the same task: look at pictures of medical scans and describe what they see. In the world of artificial intelligence, there is a popular idea called the "Platonic Representation Hypothesis." Think of it like this: if you train enough robots on enough data, they should all eventually start thinking in the exact same way. They should all build an identical internal map of reality, where a "broken bone" looks the same to Robot A as it does to Robot B, no matter who built them or what specific lessons they learned. This idea is exciting because if all medical AI models converged on this single, perfect map, doctors could swap them out like batteries without breaking their tools. But is this actually happening? Do these digital brains really agree on what a disease looks like, or are they just speaking different dialects of the same language?
A team of researchers decided to put this idea to the test with a massive, controlled experiment. They gathered 18 different "image encoders" (the parts of AI that look at pictures) and 7 "text encoders" (the parts that read reports) from various labs and companies. These models ranged from tiny ones with 7 million parameters to giants with 27 billion parameters. The researchers didn't just compare them; they dissected them. They trained new models that were identical in every way except for one thing: the goal they were trying to achieve. Some were trained with "self-supervision" (learning patterns from the images alone), some with "clinical supervision" (learning from doctors' labels), and some with "image-text" pairing (matching pictures to reports).
Here is the twist: the researchers found that the idea of a universal, shared map is mostly a myth, but a very specific kind of reality exists instead. They discovered that self-supervised learning (learning from the images themselves) is the true glue that makes these models agree with each other. In fact, adding a doctor's label to the training actually made the models less similar to their peers, not more. The models trained to just "look and learn" without specific labels ended up aligning much better (about 40% similarity in chest X-rays) than those trained with specific medical labels (only about 21%).
However, this agreement is modest and has strict limits. It only happens within the same type of picture (X-rays agree with X-rays, but not with skin photos). It doesn't get better just because the model gets bigger or smarter; a giant 27-billion-parameter model isn't necessarily more aligned than a smaller one. Most surprisingly, these image models and the text models that read medical reports do not share a common map at all. Their internal languages are completely different, and as you feed them more data, they actually drift further apart.
Despite these limitations, there is a silver lining. Even though the models don't agree on a perfect, shared geometry, they are "good enough" to be interchangeable for practical tasks. The researchers showed that a diagnostic tool trained on one model could be transferred to a completely different model and still work with about 85% of its original accuracy. It's like having two different maps of a city that don't look exactly the same, but if you know how to translate the landmarks, you can still drive from point A to point B without getting lost. The study concludes that we shouldn't wait for AI models to magically converge into one perfect brain through sheer size; instead, we need to design them with the right training goals to ensure they can at least talk to each other when it matters most.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.