← Latest papers
💻 computer science

LUX: A Lesion-Aware Graph-Conditioned Visual - Language Architecture for Explainable Endoscopic Captioning

LUX is a novel graph-conditioned vision-language architecture that enhances explainable endoscopic captioning for ulcerative colitis by constructing lesion-centric scene graphs to guide token-level generation, thereby improving clinical accuracy, interpretability, and grounding while reducing hallucinations compared to existing models.

Original authors: Alexis Ivan Escamilla-Lopez, Gilberto Ochoa-Ruiz, Salvador Hinojosa, Sharib Ali

Published 2026-08-26
📖 5 min read🧠 Deep dive

Original authors: Alexis Ivan Escamilla-Lopez, Gilberto Ochoa-Ruiz, Salvador Hinojosa, Sharib Ali

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Inside the human colon, a chronic condition known as ulcerative colitis causes the lining to become inflamed, red, and fragile. To understand how severe this inflammation is, doctors rely on a procedure called an endoscopy, where a flexible camera is passed through the digestive tract to capture images of the tissue. A specialist then examines these pictures, looking for specific signs like bleeding, ulcers, or a loss of the normal blood vessel pattern. Based on what they see, they assign a severity score and write a detailed description of the disease's state. This process is vital for deciding on treatment, but it is also difficult and subjective. Two different doctors might look at the same image and disagree on how sick the patient is, or one might miss a subtle sign of trouble that another would catch.

For years, scientists have tried to build computer programs that can read these medical images and write their own descriptions, hoping to help doctors be more consistent and accurate. Early attempts used artificial intelligence to look at the whole picture at once, compressing the entire image into a single summary before trying to write a sentence. While these systems could produce fluent text, they often failed to connect their words to the actual spots of disease in the image. They might describe bleeding that wasn't there or miss a severe ulcer because they were looking at the "big picture" rather than the specific details. This lack of connection made the computer's reports unreliable for clinical use, as the machine could not explain why it wrote what it did.

A team of researchers has now developed a new system called LUX that changes how computers approach this task. Instead of treating an endoscopic image as a single, blurry whole, LUX breaks the image down into specific, meaningful parts. When the system looks at a picture of an inflamed colon, it first identifies the exact locations of the problem areas, such as a patch of redness or a small sore. It then treats these spots as individual characters in a story, mapping out how they relate to one another. If a sore is next to a bleeding area, the system notes that connection. It builds a structured map of the disease, where every node represents a specific lesion and every line between them represents a relationship, like proximity or similarity.

This map of lesions becomes the foundation for the computer's writing. Rather than guessing what to say based on a general impression of the image, the system uses this map to guide every word it generates. As it constructs a sentence, it constantly checks its work against the specific spots it has identified. If it writes the word "bleeding," the system ensures that this word is directly tied to a specific area in the image where bleeding was detected. This method forces the computer to ground its language in visual evidence, preventing it from inventing symptoms that do not exist. The result is a description that not only sounds professional but also points directly to the physical evidence supporting each claim.

The researchers tested this approach against many other methods, including systems based on massive, general-purpose artificial intelligence models that have read millions of documents. They found that LUX was significantly better at describing the images accurately. It produced fewer errors, such as describing a healthy colon as diseased or missing a severe ulcer. In tests measuring how well the computer's descriptions matched those written by human experts, LUX outperformed all other systems. It was particularly successful at reducing "hallucinations," a term researchers use for when a computer confidently describes something that is not actually there. While other systems made these mistakes about 9.4 percent of the time, LUX reduced that rate to just 5.3 percent.

The study also showed that this new method helps the computer understand the severity of the disease more accurately. In ulcerative colitis, the severity is often determined by how different signs appear together and where they are located. Because LUX maps out the relationships between these signs, it can tell the difference between a mild case and a severe one with greater precision than systems that only look at the image as a whole. When human doctors reviewed the descriptions generated by LUX, they rated them as highly accurate and clinically useful, noting that the system correctly identified specific features like vascular patterns and tissue texture.

This work suggests that for medical imaging, especially in complex fields like endoscopy, simply making a computer smarter or giving it more data is not enough. The system needs to be taught to look at the image the way a doctor does: by focusing on specific, localized problems and understanding how they fit together. By building a structured map of the disease before it tries to write a sentence, LUX bridges the gap between seeing an image and understanding what it means. This approach offers a path toward artificial intelligence that is not just fluent, but also trustworthy and transparent, providing doctors with a tool that can explain its reasoning just as clearly as it describes the patient's condition.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →