← Latest papers
💻 bioinformatics

POME: Graph-based embeddings for partially observed mixed-type data

The paper introduces POME, a self-supervised graph-based embedding model designed to generate low-dimensional representations for partially observed mixed-type biomedical data, achieving state-of-the-art imputation performance and enabling effective downstream tasks such as patient subgroup discovery, predictive modeling, and therapy recommendation.

Original authors: Woller, F., Arend, L., Kist, A. M., List, M., Rahimi, F., Sirocchi, C., Blumenthal, D. B.

Published 2026-09-15
📖 4 min read☕ Coffee break read

Original authors: Woller, F., Arend, L., Kist, A. M., List, M., Rahimi, F., Sirocchi, C., Blumenthal, D. B.

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). ⚕️ This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer

In the vast landscape of modern medicine, data is the primary tool for understanding disease and guiding treatment. Doctors and researchers collect immense amounts of information about patients, ranging from the specific types of cells found in a tumor to the results of blood tests and the history of past illnesses. This information is rarely uniform; it is a chaotic mix of numbers, categories, and text, often referred to as mixed-type data. Furthermore, this data is almost never complete. Patients miss appointments, certain tests are not ordered for everyone, and records are lost, leaving behind gaps that look like empty spaces in a spreadsheet. For decades, these gaps and the messy variety of data types have forced scientists to either throw away incomplete records or use statistical tricks to fill them in, both of which can distort the true picture of a patient's health and lead to flawed conclusions.

A team of researchers has now developed a new way to handle this messy reality, turning the gaps and the variety into features rather than flaws. They created a tool called POME, which stands for partially observed mixed-type data embeddings. Instead of trying to force all the different kinds of medical data into a single, rigid format or simply deleting the missing pieces, POME treats the entire dataset as a map of connections. Imagine a network where patients are one set of points and every possible piece of information about them is another set of points. When a patient has a blood test result, a line connects them to that result. When a piece of data is missing, the line simply isn't drawn. This approach allows the computer to learn the relationships between patients and their symptoms, even when the data is incomplete and jumbled together.

The researchers tested this approach on three large, real-world collections of medical records involving thousands of patients with lung cancer, head and neck cancer, and various other conditions treated with chemotherapy. They found that POME could fill in the missing information with a level of accuracy that matched or exceeded the best existing methods. More importantly, the tool created a new kind of "fingerprint" for each patient. These fingerprints were not just simple summaries but rich, low-dimensional representations that captured the complex interplay of a patient's diverse health markers. When the researchers used these fingerprints to group patients together without telling the computer what to look for, the groups that formed were medically meaningful. For instance, the tool naturally separated patients based on whether their tumors were linked to a specific virus, a factor known to influence survival rates, or grouped them by the density of immune cells in their tumors, which is a sign of how the body is fighting the disease.

The power of these fingerprints extended beyond just sorting patients. The researchers used them to look back at historical treatment decisions and see if they made sense. By comparing a new patient's fingerprint to those of past patients, the system could suggest which treatment strategy had worked best for similar individuals in the past. In one test involving head and neck cancer patients, those whose actual treatment matched the system's recommendation based on these fingerprints had significantly better five-year survival rates than those whose treatment did not match. This suggests that the tool can identify subtle patterns in how different combinations of symptoms and history respond to specific therapies, offering a way to support doctors in making more informed choices.

The study also showed that these fingerprints could be used to predict future outcomes, such as whether a patient would develop a specific side effect from chemotherapy, often outperforming other standard methods for handling mixed data. The researchers even looked inside the tool to see what it had learned about the variables themselves. They found that the tool had correctly grouped together medical terms that are biologically related, such as different types of white blood cells, demonstrating that it had learned the underlying logic of the medical data rather than just memorizing numbers.

While the tool is currently optimized to run on powerful computer hardware and requires significant memory for very large datasets, the researchers demonstrated that it is fast enough to be useful in real-world workflows. They made the software available to the public so that other scientists can test it on their own data. The work does not claim to solve every problem in medical data analysis, but it offers a robust new way to turn the messy, incomplete, and varied records of clinical practice into clear, actionable insights. By respecting the natural structure of the data, including its missing parts, POME allows researchers to see the patient more clearly, potentially leading to better diagnoses and more effective treatments.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →