← Latest papers
🧬 biology

Efficiency and Interpretability in MYC Status Prediction: A Comparative Study of Vision Transformers

This study benchmarks five Vision Transformer architectures for predicting MYC status in Diffuse Large B-cell Lymphoma, revealing that while the Swin Transformer offers superior speed and risk stratification, the Hierarchical ViT achieves the highest diagnostic accuracy, thereby establishing a mathematically validated framework for optimizing architectural trade-offs in digital pathology.

Original authors: GEI KI TANG, CHEE CHIN LIM, FAEZAHTUL ARBAEYAH HUSSAIN, QI WEI OUNG, AIDY IRMAN YAJID, SUMAYYAH MOHAMMAD AZMI

Published 2026-09-02
📖 4 min read☕ Coffee break read

Original authors: GEI KI TANG, CHEE CHIN LIM, FAEZAHTUL ARBAEYAH HUSSAIN, QI WEI OUNG, AIDY IRMAN YAJID, SUMAYYAH MOHAMMAD AZMI

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). ⚕️ This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer

In the world of cancer care, speed and precision are often locked in a difficult struggle. Doctors need to know exactly what kind of disease they are fighting to choose the right treatment, but the tests that provide this certainty can be slow, expensive, and require specialized equipment. One such critical test involves looking for a specific genetic switch called MYC in a type of aggressive blood cancer known as diffuse large B-cell lymphoma. When this switch is turned on, the disease tends to grow faster and respond poorly to standard treatments, making its early detection vital for patient survival. Currently, the gold standard for finding this switch is a laboratory technique that uses glowing markers to highlight the genetic material under a microscope. While accurate, this process takes time and resources that are not always available in every clinic. This creates a pressing need for a faster way to screen patients, one that could look at standard microscope slides and instantly flag those who need the more intensive, expensive testing.

A team of researchers from Malaysia has taken a significant step toward solving this problem by teaching computers to read these standard microscope slides. Instead of using the traditional tools of artificial intelligence, which often struggle to see the big picture in complex biological images, the team tested a newer generation of models called Vision Transformers. These models are designed to look at an image in pieces and understand how those pieces relate to one another across the entire picture, much like a pathologist scanning a slide to see both individual cells and the overall pattern of the tissue. The researchers pitted five different versions of this technology against each other to see which one could best predict the presence of the MYC gene. They fed the models hundreds of thousands of tiny, high-resolution snapshots taken from digital slides of patient tissue, asking the computers to decide if each snapshot came from a patient with the aggressive genetic marker or without it.

The results revealed a clear trade-off between how fast a model works and how accurately it diagnoses the condition. One model, known as the Hierarchical Vision Transformer, proved to be the most accurate. It correctly identified the genetic status in the vast majority of cases, matching the performance of a seasoned expert. This model works by mimicking the way a human pathologist examines a slide, zooming in to see fine details of individual cells while also stepping back to understand the broader landscape of the tissue. However, this deep level of analysis comes at a cost: it is the slowest of the group, taking the most time to process each image. On the other hand, a model called the Swin Transformer offered a different balance. It was nearly three times faster than the most accurate model, processing images with remarkable speed, though it made slightly more mistakes in its final yes-or-no decisions.

Despite the difference in speed, both top-performing models showed they were looking at the right things. When the researchers used a technique to visualize what the computers were focusing on, the images revealed that the models were zeroing in on the specific shapes and textures of the cancer cells, such as the size and color of the cell nuclei, rather than getting distracted by empty space or background noise. This is a crucial finding because it suggests the models are learning the actual biological signs of the disease, not just guessing based on random patterns. The study also highlighted that while some older, simpler models were faster, they failed to capture the complex details needed for this specific task, and models trained on general images from the internet did not perform well enough without specific medical training.

The researchers concluded that there is no single perfect tool for every situation, but rather two excellent options depending on the goal. If a hospital needs the absolute highest level of diagnostic certainty to confirm a difficult case, the slower, more detailed Hierarchical model is the best choice. If the goal is to quickly screen a large number of patients to flag those who might need further testing, the faster Swin Transformer is highly effective. The study emphasizes that while these computer systems are powerful, they are currently designed to work on small pieces of a slide and would need further testing across different hospitals and with whole-slide analysis before they could be used routinely in clinics. For now, this work provides a clear blueprint for how artificial intelligence can be tailored to handle the complex, multi-layered nature of cancer tissue, offering a promising path toward faster and more accessible cancer care.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →