← Latest papers
📄 medicine

Nonparametric Information Sufficiency: Nonidentifiability, Deformation, and the Limits of Prediction from Observed Biomedical States

This study demonstrates that for the vast majority of observed patient states across 13 clinical substrates, the available information is fundamentally insufficient to distinguish between divergent future health trajectories, establishing nonidentifiability as an intrinsic data limitation rather than a failure of predictive model selection.

Original authors: Maurice Antony Ewing

Published 2026-09-03
📖 8 min read🧠 Deep dive

Original authors: Maurice Antony Ewing

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

In the world of modern medicine, doctors and computer systems often rely on a simple assumption: if two patients look the same right now, they will likely follow the same path forward. If two people have the same blood pressure, the same tumor size, and the same symptoms, the expectation is that they will respond to treatment similarly or that their diseases will progress at the same rate. This belief underpins much of how we predict health outcomes and design personalized care plans. However, this assumption rests on the idea that the information we can see and measure at a single moment is enough to tell the whole story of a person's future health. It assumes that the current snapshot contains all the necessary clues to predict what comes next. But what if the snapshot is incomplete? What if two people standing in the exact same spot on a map are actually walking toward completely different destinations, not because they are confused, but because the map itself is missing the terrain that separates their paths?

This is the central question explored in a new study by Maurice Antony Ewing, a researcher at the University of Illinois at Chicago. The work challenges the idea that better computer models or more complex algorithms can solve the problem of predicting patient outcomes. Instead, the study suggests that the problem lies not in the tools we use to analyze data, but in the data itself. The researchers investigated whether the information available at the moment a doctor makes a decision is actually sufficient to distinguish between different possible futures. They found that for a vast number of patients, the answer is no. Two people can appear identical based on every measurement a doctor can take, yet one might get better while the other gets worse. The study argues that this is not a failure of prediction technology, but a fundamental limit of the information available.

To understand this, imagine two patients with a specific disease. At the moment a doctor checks them, their vital signs, lab results, and imaging scans are nearly identical. A standard prediction model would look at these numbers and assign them the same likely outcome. But in reality, one patient might recover quickly while the other deteriorates. The study asks: why does this happen? Is it because the computer model is not smart enough to find a hidden pattern? Or is it because the pattern simply does not exist in the numbers the doctor has? The researchers call this phenomenon "non-identifiability." It means that the current state of the patient does not contain enough unique information to identify which future path belongs to them. The future is not just uncertain; it is fundamentally unidentifiable from the present observations alone.

To test this, the researchers gathered a massive amount of real-world medical data. They looked at 13 different human clinical datasets, covering a wide range of conditions including neurodegenerative diseases like Alzheimer's, critical care scenarios, cancer, and heart disease. In total, they examined nearly 700,000 individual health observations recorded over time. They focused on five large groups of patients who had been tracked for a long time, involving more than 56,000 individuals and over 400,000 specific moments in their health journeys. For each moment in a patient's history, the researchers asked a simple question: if we find other patients who look exactly like this one at this specific moment, do they all go on to have the same future?

The results were striking. In the five major groups studied, the researchers found that in nearly every case, patients who looked the same at a specific moment went on to have very different futures. Specifically, between 65% and 99% of the patient states they examined were "non-identifiable." This means that for the vast majority of these moments, the information available did not distinguish between a future of improvement and a future of decline. In some groups, like those with Parkinson's disease or amyotrophic lateral sclerosis, the number was as high as 98.7%. Even more telling, the researchers looked for cases where patients with similar states went in opposite directions—one getting better and the other getting worse. They found that in between 59% and 82% of these matched groups, both positive and negative outcomes were happening. This proves that the ambiguity is not just about whether a patient stays the same or changes; it is about the fact that similar patients are actively moving in opposite directions.

The study also addressed a common counter-argument: perhaps the data is there, but the computer models are just not sophisticated enough to find it. To test this, the researchers did not rely on a single type of artificial intelligence. Instead, they used a "zoo" of different models, ranging from simple statistical tools to complex deep learning systems. They tested these models across different types of data and different disease scenarios. The result was consistent: no matter how advanced the model was, it could not predict the future direction of these patients any better than the others. When the models failed, they all failed in the same way. This suggests that the limitation is not in the architecture of the computer program, but in the information fed into it. A model cannot invent a distinction that is not present in the data. If the observed state does not contain the clue that separates a recovering patient from a declining one, no amount of mathematical complexity can create that clue.

The researchers also looked at whether adding more history would solve the problem. Intuitively, one might think that knowing a patient's past trajectory—how they got to their current state—would help predict where they are going next. The study found that while history sometimes helps, it often does not. In many cases, even the full history of a patient's condition was compatible with multiple, opposing futures. Two patients might have had very different paths to reach the same point, yet their subsequent paths diverged in ways that their histories did not predict. This means that simply collecting more rows of data or looking further back in time does not automatically solve the problem. The key is not the quantity of observations, but whether those observations contain the specific information needed to separate the different possible futures.

This finding has profound implications for how we approach medical prediction. For years, the focus in medical artificial intelligence has been on building better models, finding more data, and improving algorithms. The study suggests that this approach hits a hard ceiling. If the information required to distinguish between two patients is missing from the clinical record, then no model can bridge that gap. The problem is not that we are asking the wrong question of the data; it is that the data itself is insufficient to answer the question. The researchers emphasize that this is not a claim that patients are inherently unpredictable in a mystical sense, but rather that the specific information we choose to measure and record is often not enough to tell the story of their future.

The study does not claim that we should stop trying to predict outcomes or that we cannot improve care. Instead, it points to a different path forward. If the current measurements are not enough, then the solution lies in finding new kinds of information. This could mean measuring different biological markers, using different types of imaging, or looking at factors that are currently ignored. It could also mean accepting that for some patients, at some moments, the available information simply does not allow for a precise prediction, and that the best course of action might be to gather more specific data before making a decision. The study concludes that before we spend more resources on building smarter models, we must first ask whether the information we have is actually sufficient to distinguish the futures we are trying to predict.

In the end, the research offers a sobering but necessary reality check for the field of precision medicine. It shows that the gap between what we can see and what will happen is often wider than we thought. Two patients can stand side by side, looking identical in every way that modern medicine can measure, yet be on entirely different trajectories. The tools we use to predict their futures are not broken; they are simply working with incomplete maps. The path forward requires us to look beyond the current state of the art in modeling and focus on the fundamental question of what information is actually needed to tell one patient's story from another's. Until we find that missing information, the limits of prediction will remain, not because of a lack of computing power, but because of a lack of distinguishing detail in the data itself.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →