A Unified DINOv2-Based Framework for LVEF Estimation, GLS Dysfunction Classification, and Early Cardiotoxicity Prediction
This paper presents a unified DINOv2-based framework that leverages frozen foundation models with task-specific adaptations to achieve robust, cycle-free estimation of left ventricular ejection fraction, classification of global longitudinal strain dysfunction, and early cardiotoxicity prediction, while further enhancing LVEF accuracy through a specialized physiology-guided hybrid regression model.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
The heart is a tireless pump, but like any machine, it can wear down, especially when fighting cancer. Chemotherapy drugs are powerful weapons against tumors, yet they sometimes damage the heart muscle, a side effect known as cardiotoxicity. Detecting this damage early is a race against time. Doctors currently rely on two main ways to check the heart's health. The first is a measure of how much blood the heart pushes out with each beat, a standard metric that tells the big picture. The second is a more sensitive tool that tracks the subtle stretching and squeezing of the heart muscle itself, often revealing trouble before the big picture shows any sign of strain. The challenge lies in the fact that these signals are hidden inside complex, moving images of the heart, and spotting the earliest warning signs often requires a human expert to manually trace the heart's boundaries and count its beats, a process that is slow and prone to human error.
A team of researchers at GE HealthCare has developed a new computer system designed to read these heart videos with greater speed and accuracy, aiming to spot trouble before it becomes irreversible. Instead of teaching a computer from scratch to understand heart images, they built their system on top of a massive, pre-trained "foundation" model. Think of this foundation model as a student who has already studied millions of general images and learned to recognize shapes, textures, and patterns. The researchers then gave this student a specialized course on heart videos, teaching it to focus on the specific details needed to measure heart function and predict drug-related damage. This approach allowed the system to learn quickly and accurately, even with a relatively small number of patient videos available for training.
The researchers tested their system on three distinct tasks using a large collection of heart ultrasound videos from patients across Europe. The first task was to calculate the percentage of blood the heart pumps out, a critical number for diagnosing heart failure. The second was to classify whether the heart muscle was showing signs of dysfunction based on its stretching patterns. The third, and perhaps most difficult, was to predict which patients would develop heart damage from chemotherapy before they even started treatment, using only their baseline heart scans. To handle these different goals, the system used a clever strategy where it kept its core knowledge frozen but added small, flexible "adapters" for each specific job. This meant the system could learn the unique visual cues for pumping volume without confusing them with the cues for muscle stretching, ensuring that each task was performed with high precision.
A key innovation in this work was how the system handled the timing of the heart's beat. In a traditional approach, a computer might need to be told exactly when the heart is fully relaxed and when it is fully squeezed to make a measurement. This new system, however, learned to understand the heart's rhythm on its own. It could infer the necessary timing information directly from the visual content of the video, without needing any external labels or manual markings to tell it where the start and end of a beat were. This made the system much more robust and practical for real-world use, where such precise manual labels are often missing. For the specific task of measuring pumping volume, the team also created a separate, specialized model that combined the full video motion with a focused look at the heart's most critical moments of relaxation and contraction, mimicking how a human doctor might mentally compare the heart's state at different times.
The results of the study showed that this new approach worked remarkably well. For measuring the heart's pumping volume, the system achieved an average error of just over five percent, and the specialized model improved this to under five percent, a level of accuracy that rivals human experts. When it came to detecting subtle muscle dysfunction, the system correctly identified the problem in nearly 77 percent of cases. Most notably, for predicting early heart damage from chemotherapy, the system's ability to spot at-risk patients improved significantly, correctly ranking patients by risk in over 70 percent of cases, a substantial jump from previous methods. These findings suggest that by leveraging powerful pre-trained visual knowledge and adapting it carefully to medical needs, computers can become powerful allies in protecting cancer patients from heart damage, offering a way to intervene early and potentially save lives.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.