Braided Vision Transformer for Stroke Detection in Multi-view Retinal Fundus Imaging
This paper introduces the Braided Vision Transformer (BViT), a novel deep learning model that leverages multi-view retinal fundus images from both eyes to detect stroke and transient ischemic attacks by simultaneously extracting features and modeling inter-view relationships, achieving an AUC of 0.75 on a newly collected dataset.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Stroke remains one of the world's most devastating health challenges, a leading cause of death and long-term disability that strikes without warning. For decades, doctors have relied on heavy, expensive machines like CT scanners and MRIs to see inside the brain and confirm a stroke. These tools are essential for treatment, but they are not always available for quick screening, especially in remote areas or for early warning signs. A growing body of research suggests that the key to spotting these dangers might lie not in the brain itself, but in the eye. The retina, the light-sensitive tissue at the back of the eye, is the only place in the body where doctors can directly observe blood vessels without surgery. Because the vessels in the eye and the brain are so similar, changes in the tiny blood vessels of the retina often mirror the health of the brain's own circulatory system. This connection has sparked a search for a faster, cheaper, and non-invasive way to assess stroke risk using nothing more than a photograph of the eye.
Researchers at the VTT Technical Research Centre of Finland have taken this concept a step further by developing a new computer system designed to read these retinal photographs with unprecedented depth. In a recent study, they introduced a method called the Braided Vision Transformer, a type of artificial intelligence that looks at the eye in a way no previous system has. Instead of analyzing a single image in isolation, this system examines four distinct views of the retina at once: the macula and the optic nerve head from both the left and right eyes. The macula is the central part of the retina responsible for sharp vision, while the optic nerve head is where the nerve fibers exit the eye to travel to the brain. By capturing these specific areas from both eyes, the system gathers a much richer picture of a person's vascular health than a single snapshot ever could.
The core innovation of this work lies in how the computer processes these four images. Traditional artificial intelligence models often treat each image as a separate puzzle piece, solving them one by one before combining the answers. The new Braided Vision Transformer, however, weaves the information together from the very beginning. It creates a structure where the computer constantly compares the blood vessel patterns in the left eye against those in the right eye, and the central view against the side view, all at the same time. This approach mimics how a skilled clinician might look at a patient, constantly shifting focus between different parts of the eye and both eyes to spot subtle inconsistencies that signal trouble. The researchers trained this system on a dataset of over 800 retinal images collected from more than 200 participants, including people who had suffered a stroke, those who had experienced a transient ischemic attack—a temporary blockage often called a mini-stroke—and healthy individuals.
The results of this training were promising. When tested on its ability to distinguish between healthy eyes and those showing signs of stroke or mini-stroke, the new system achieved a performance score of 0.75, a measure of accuracy that was higher than other standard computer vision models tested in the same study. The system proved particularly effective at recognizing the complex patterns associated with cerebrovascular events, outperforming models that looked at images individually or used older methods of combining data. The study also highlighted a crucial finding: looking at multiple views of the eye is significantly better than looking at just one. When the researchers tested a version of the system that only looked at a single image, its ability to detect stroke dropped dramatically, suggesting that the subtle signs of vascular disease are distributed across different parts of the retina and both eyes, and are easily missed if the view is too narrow.
While the system is not yet ready to replace the CT scanner in a hospital emergency room, it represents a significant leap forward in the potential for early screening. The researchers note that their approach offers a portable and cost-effective alternative that could one day be used in community clinics or even remote locations to flag individuals who need immediate, more detailed medical attention. The study confirms that the retina holds valuable clues about brain health and that modern artificial intelligence, when designed to look at the whole picture rather than just a part, can unlock those clues more effectively than before. By successfully weaving together multiple views of the eye, this new method suggests a future where the first line of defense against stroke might be as simple as a quick, painless photograph.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.