← Latest papers
💻 computer science

A Hybrid CNN-Transformer Framework for Robust Coronary Artery Stenosis Detection in X-ray Angiography

This study proposes a hybrid CNN-Transformer framework based on RT-DETR that achieves state-of-the-art accuracy (97.8% mAP50) and real-time performance (38 FPS) for robust coronary artery stenosis detection in X-ray angiography, significantly outperforming traditional methods in both precision and speed.

Original authors: Hossein Sadr, Kamran Balani, Ali. A. Kiaie, Zeynab Khodaverdian, Arsalan Salari, Mahboobeh Hoseinalizadeh, Mojdeh Nazari

Published 2026-09-21
📖 5 min read🧠 Deep dive

Original authors: Hossein Sadr, Kamran Balani, Ali. A. Kiaie, Zeynab Khodaverdian, Arsalan Salari, Mahboobeh Hoseinalizadeh, Mojdeh Nazari

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). ✨ This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Heart disease remains the leading cause of death worldwide, and at the heart of the problem is often a narrowing of the arteries that supply blood to the heart muscle. This condition, known as coronary artery stenosis, occurs when fatty deposits and calcium build up inside the vessel walls, restricting the flow of oxygen-rich blood. If this narrowing goes undetected or untreated, it can lead to a heart attack. Doctors currently rely on X-ray angiography to see inside these arteries. In this procedure, a contrast dye is injected into the bloodstream, making the vessels visible on an X-ray screen. However, reading these images is difficult. The pictures are often grainy, filled with shadows from bones like ribs and vertebrae, and cluttered with medical tools. Interpreting them requires a human expert to look closely at every frame, a process that is slow, tiring, and prone to human error.

To help doctors, scientists have turned to artificial intelligence, specifically a type of computer program called deep learning. For years, the most successful programs for this task were built on a technology called Convolutional Neural Networks, or CNNs. These programs are excellent at spotting local details, such as the sharp edge of a vessel wall or a specific texture. However, they struggle to understand the bigger picture. Because they focus so intently on small, nearby pixels, they often get confused by the noisy background, mistaking a rib or a catheter for a blocked artery. They lack the ability to step back and see the entire path of the blood vessel to understand its true shape. A newer technology, known as the Transformer, solves this by allowing the computer to look at the whole image at once, weighing the importance of every part relative to every other part. The challenge has been combining the speed of the older technology with the global understanding of the new one.

In this study, researchers from several medical universities in Iran proposed a new solution that merges these two approaches. They developed a hybrid framework that uses a Real-Time Detection Transformer, or RT-DETR, to find blocked arteries in X-ray images. Instead of relying on a single type of computer vision, their system uses a standard network to grab the fine details of the image and then passes that information to a Transformer module. This second module acts like a global map, helping the computer distinguish between a real blockage and a confusing shadow cast by a rib or a spine. The goal was to create a tool that is not only highly accurate but also fast enough to be used in real-time during a medical procedure.

The team trained their new system using a large collection of 8,325 X-ray images taken from 100 different patients. These images came from a research institute in Russia and included a variety of real-world conditions, with different resolutions and levels of noise. To teach the computer how to handle these difficult images, the researchers used a technique where they stitched four different images together into one large training picture. This forced the system to learn how to spot small blockages even when they were surrounded by complex and messy backgrounds. They also used another method that blended pairs of images together to create new, virtual examples, helping the computer become more robust against the random noise found in medical scans.

When the researchers tested their new framework against existing methods, the results were clear. The hybrid model achieved a detection accuracy of 97.8 percent, which is higher than the previous best models that relied solely on older technology. One of the most significant improvements was in finding small blockages. The new system correctly identified 59.2 percent of these tiny lesions, a substantial jump compared to the 51.5 percent achieved by the leading previous model. This matters because small blockages are often the hardest to see but can be the most dangerous if missed. The researchers also found that their system was remarkably fast, processing images at a rate of 38 frames per second. This is more than three times faster than the previous top-performing model, which managed only 11 frames per second.

The study suggests that this new approach successfully bridges the gap between high accuracy and the speed required for clinical use. By combining the ability to see fine details with the power to understand the entire image context, the system reduces the number of false alarms caused by background noise. The authors note that while the results are promising, the system was tested on data from a single source, and further work is needed to ensure it works equally well on images from different hospitals and machines. Nevertheless, the findings indicate that this hybrid framework offers a robust and efficient tool that could one day assist cardiologists in making faster, more reliable decisions during critical heart procedures.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →