← Latest papers
🧬 biology

Application of H&E images combined with gene expression data-driven multimodal deep learning models in gastric cancer prognosis

This study introduces PGC-pro, a multimodal deep learning model that integrates H&E histopathology images, gene expression profiles, and clinical data to significantly improve gastric cancer prognostic accuracy and interpretability compared to single-modal approaches.

Original authors: Bo Yang, Hongyi Cai, Hao Li

Published 2026-09-07
📖 6 min read🧠 Deep dive

Original authors: Bo Yang, Hongyi Cai, Hao Li

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). ⚕️ This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer

Gastric cancer, often called stomach cancer, remains one of the most formidable health challenges worldwide, claiming hundreds of thousands of lives each year. For decades, doctors have relied on a system called TNM staging to predict how a patient will fare. This system acts like a map, charting the size of a tumor and how far it has spread to nearby lymph nodes or distant organs. While this map is essential, it is also incomplete. It describes the physical extent of the disease but often misses the unique biological personality of the cancer itself. Two patients with tumors of the same size and location can have vastly different outcomes because their cancers behave differently at a microscopic level. To bridge this gap, scientists are turning to two powerful sources of information that have traditionally been kept separate: the visual patterns seen under a microscope and the molecular instructions written inside the cells. One source is the H&E slide, a standard glass slide where a thin slice of tissue is stained with pink and purple dyes to reveal the shape and arrangement of cells. The other is gene expression data, a digital readout of which genes are active within the tumor, acting as a molecular fingerprint. The challenge has been to combine these two very different types of information—images and numbers—into a single, coherent picture that can tell a doctor exactly how dangerous a specific patient's cancer is likely to be.

In a recent study, researchers from Qinghai University Affiliated Hospital tackled this challenge by building a new kind of computer model designed to learn from both the visual and molecular worlds simultaneously. They called their creation PGC-pro, a system that does not just look at a tumor's shape or its genes in isolation, but studies how they talk to each other. The team started with data from 318 patients who had been treated for stomach cancer. For each patient, they gathered three things: a digital scan of the entire tumor slide, a list of how active thousands of genes were in that tumor, and the patient's clinical history, including age and the stage of the disease. The researchers first taught the computer to recognize the most important parts of the microscopic images. Using a method that mimics how a pathologist scans a slide, the system learned to ignore empty spaces and focus on the crowded areas where cancer cells are invading normal tissue. Simultaneously, the computer sifted through the genetic data to find the specific genes that were most strongly linked to how long patients survived, narrowing thousands of possibilities down to a core group of 150 genes that truly mattered.

The core innovation of this work lies in how the computer connects these two streams of information. Instead of simply stacking the image data next to the gene data, the model uses a mechanism that allows the two to ask questions of one another. Imagine the image data asking the gene data, "What molecular changes are happening in this specific cluster of cells?" while the gene data asks the image data, "What does the tissue look like where these genes are most active?" This back-and-forth exchange allows the model to find deep connections that neither side could see alone. For instance, the model might learn that a specific messy arrangement of cells in the image is almost always paired with a high level of activity in certain genes, and that this specific combination is a strong warning sign for a poor outcome. By integrating these visual and molecular clues with the patient's age and disease stage, the model generates a single risk score for each person.

The results of this approach were clear and measurable. When the researchers tested the model on the patient data, it proved significantly better at predicting survival than any method that used only one type of information. A model looking only at the genes could predict outcomes with a certain level of accuracy, and a model looking only at the images did slightly worse, but the combined model outperformed both. The system achieved a score that indicates a high ability to correctly rank patients by their risk of death, a figure that rose noticeably when all three data sources were used together. More importantly, the model did not act as a mysterious "black box" that gives an answer without explanation. The researchers built in tools that let them see exactly what the computer was looking at. They could generate heat maps showing which parts of the tumor slide the model focused on, and these maps consistently highlighted the edges of the tumor where it was invading healthy tissue, as well as areas of dead cells. They could also see which genes were driving the prediction, identifying specific molecular markers that the model had learned were critical. This transparency is vital, as it allows doctors to trust the computer's judgment by seeing the biological evidence behind it.

The study also tested how well the model would hold up if it had to make predictions over time. It successfully estimated survival probabilities at one, three, and five years, maintaining its accuracy across these different timeframes. When the patients were split into high-risk and low-risk groups based on the model's scores, the difference in their actual survival rates was stark and statistically significant, confirming that the model could effectively separate those who would likely do well from those who would not. The researchers noted that while the model performed well, it was trained on a single database of patient records, and its ability to work on patients from different hospitals or regions still needs to be tested. They also acknowledged that the images they used sometimes had imperfections, like uneven staining, which could affect the computer's view. Despite these limitations, the work demonstrates that combining the visual story of a tumor with its molecular story creates a much more powerful tool for understanding the disease. By proving that these different types of data can be woven together to reveal a clearer picture of the future, this research offers a promising path toward more personalized and accurate care for patients facing gastric cancer.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →