← Latest papers
📄 medicine

Real-time detection of High-Risk Laryngeal Lesions using Laryngoscopic Image-Based AI-Assisted Triage: A Comparative Study of Multiple Vision Foundation Models

This study demonstrates that an AI model based on the DINOv2 vision foundation model achieves high accuracy in stratifying high-risk laryngeal lesions from laryngoscopic images and significantly improves diagnostic performance across physicians of varying experience levels.

Original authors: Junyi Wang, Tianjian Zhang, Xing Huang, Ying Wang, Sisi Zhang, Yu Zhou, Yunrong Ti, Yingpeng Xu, Chuanyao Lin, Xiaoyun Qian

Published 2026-08-28
📖 4 min read☕ Coffee break read

Original authors: Junyi Wang, Tianjian Zhang, Xing Huang, Ying Wang, Sisi Zhang, Yu Zhou, Yunrong Ti, Yingpeng Xu, Chuanyao Lin, Xiaoyun Qian

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

The human voice is a fragile instrument, and the larynx, or voice box, is where that instrument is built. When the delicate tissues inside the throat develop abnormal growths, distinguishing between a harmless irritation and a dangerous cancer is one of the most critical tasks for an ear, nose, and throat specialist. This decision often relies on a flexible camera passed through the nose to view the vocal cords, a procedure known as laryngoscopy. While these images reveal the physical state of the throat, the visual signs of early cancer can look remarkably similar to benign conditions like polyps or cysts. In a busy clinic, a doctor must decide in seconds whether a patient needs an immediate biopsy, a surgical evaluation, or simply a watchful wait. This high-stakes judgment is difficult, and the outcome can vary depending on the doctor's experience, the quality of the image, and the subtle nature of the changes on the tissue surface.

To address this challenge, a team of researchers from hospitals in Nanjing and Johns Hopkins has developed a new type of artificial intelligence designed to act as a second pair of eyes for these specialists. Rather than trying to replace the doctor, the system is built to help sort patients into risk categories, flagging those with the highest likelihood of dangerous lesions. The researchers trained this computer system using thousands of real laryngoscopic images, teaching it to recognize the specific visual patterns that separate high-risk tissue from low-risk tissue. They did not just build one model; they tested eleven different types of underlying computer architectures, comparing how well each could learn from the data. They found that a specific model, which had been pre-trained on a vast library of general images before being specialized for the throat, performed the best. This "foundation model" was able to focus its attention on the exact areas of the vocal cords that mattered, ignoring the background noise and zeroing in on the texture and color of the tissue itself.

When the researchers tested this system on new images it had never seen before, the results were strikingly consistent. In a set of internal test images, the AI correctly identified high-risk cases with an accuracy of 96.0 percent. Even more importantly, when the system was tested on images from two different hospitals with different equipment and lighting conditions, it maintained its high performance, achieving accuracy rates of 96.7 percent and 95.5 percent. This suggests that the system is robust enough to work in real-world clinics, not just in the controlled environment where it was built. The researchers also explored how this technology could work with live video, which is how doctors actually examine patients. They developed a method to automatically pick the clearest frames from a video stream and focus only on the vocal cords, discarding blurry or irrelevant parts. This video-based approach improved the accuracy of the risk assessment, proving that the system could handle the dynamic nature of a live examination.

The true test of any medical tool, however, is how it helps the people who use it. To find this out, the researchers invited ten doctors of varying experience levels to review 120 images, first on their own and then with the AI's assistance. The results showed that the AI helped everyone, but it helped the least experienced doctors the most. When junior doctors used the AI, their accuracy improved by four percentage points, a significant jump that brought their performance closer to that of senior experts. The system also helped doctors make their decisions faster, reducing the time spent on each image. This indicates that the technology acts as a powerful equalizer, offering consistent, high-quality support that can help less experienced clinicians make safer, more confident decisions.

The study did not claim to have solved the problem of laryngeal cancer diagnosis, nor did it suggest that the AI should replace human judgment. Instead, the findings suggest that a carefully designed artificial intelligence can serve as a reliable triage tool, helping to ensure that high-risk patients are identified quickly and that benign cases are not subjected to unnecessary procedures. The researchers noted that while the system performed well, it was tested on a specific set of images and requires further study in larger, real-world clinical settings to confirm its long-term utility. Nevertheless, the work demonstrates that modern computer vision, when applied with care and tested rigorously, can learn to see the subtle signs of disease in the human throat, offering a new layer of support to the doctors who stand between patients and their health.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →