← Latest papers
💻 computer science

Evaluation of Clinically Steerable Retinal Image Generation from Foundation Model Latent Spaces

This paper evaluates the ability of retinal foundation models to generate controllable images that preserve clinical phenotypes, finding that while they outperform conventional methods within their own frameworks, they exhibit a significant synthetic-to-real representation gap when assessed against real-image classifiers.

Original authors: Zuzanna A. Wakefield-Skórniewska, Bartłomiej W. PapieĊ

Published 2026-08-14
📖 4 min read☕ Coffee break read

Original authors: Zuzanna A. Wakefield-Skórniewska, Bartłomiej W. PapieĊ

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to teach a robot to paint pictures of the human eye, not just to look pretty, but to tell a doctor exactly what is wrong with a patient's health. This is the world of medical imaging, where computers are learning to see patterns in photos that human eyes might miss. To do this, scientists use "foundation models," which are like super-smart students that have read millions of medical textbooks (in this case, millions of eye photos) without being told the answers. These models learn to understand the "language" of the eye, creating a hidden map of what healthy and sick eyes look like. The big question scientists are asking is: Can we use this hidden map to not just understand eyes, but to create new, fake eye photos that are so realistic and accurate they can help train other doctors or test new treatments? If we can steer these AI models to draw eyes with specific traits—like high blood pressure or a history of heart trouble—we could generate endless training data without needing real patients. But there's a catch: just because the AI thinks it's drawing a perfect sick eye, does it actually look like a real sick eye to a human doctor or a different computer program?

This paper dives into that exact challenge by testing four different "super-student" AI models designed for retinal (eye) imaging. The researchers wanted to see if they could use these models as a blueprint to generate synthetic eye images that faithfully carry specific health information, like age, sex, or the risk of future heart attacks. They used a clever two-step trick: first, they taught a generator to create the "idea" of an eye based on health data (like "this person is 60 and has high blood pressure"), and then they used a decoder to turn that idea into a full-color photo.

The results were a mix of "wow" and "wait a minute." When the researchers checked the generated images using the same AI model that created them, the results were fantastic. The fake images were like perfect carbon copies; they carried all the health information encoded in the AI's brain. If the AI was told to draw an eye for a 60-year-old with high blood pressure, the resulting image was so convincing that the AI's own internal tests said, "Yes, this is definitely a 60-year-old with high blood pressure." In fact, these foundation-model-based creations were often better at holding onto these details than standard AI painting tools.

However, the story takes a twist when the researchers stepped outside their own bubble. They took these beautiful, AI-generated eye photos and showed them to a completely different, standard computer program (a ResNet classifier) that had only ever seen real human eye photos. Suddenly, the magic faded. The fake images, which looked perfect to the AI that made them, struggled to fool the outsider. The standard program had a much harder time guessing the age or health conditions from the synthetic photos compared to real ones. It's as if the AI learned to speak a very specific dialect of "eye language" that only it understands. While the images looked realistic to the human eye, they missed subtle, real-world textures that other computers rely on to make diagnoses.

The paper concludes that while these foundation models are powerful tools for creating controllable, health-aware eye images, there is still a "translation gap" between what the AI thinks is real and what the rest of the world sees. One model, called URFound, which was trained on both eye photos and other medical scans, did the best job of bridging this gap, but even it wasn't perfect. The authors suggest that while we are getting closer to being able to generate useful medical data on demand, we still need to teach these AI models to speak a language that is compatible with the real world, not just their own internal logic. Until that gap is closed, we must be careful about assuming that a perfect-looking AI-generated medical image is ready to replace real patient data in critical medical tasks.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →