← Latest papers
💻 bioinformatics

Incomplete references leave bulk deconvolution targets non-identifiable, but identification has a computable precision price

This paper argues that incomplete reference profiles render bulk deconvolution targets non-identifiable rather than merely difficult to estimate, demonstrating that standard point estimates often lack coverage and can reverse biological conclusions, while proposing a new framework that prioritizes certified precision, coverage, and false-certification metrics over simple shrinkage.

Original authors: Jiang, H., Gao, F., Liu, P., Wu, Y., Jie, Y., Li, Y., Jiang, Y.

Published 2026-09-18
📖 6 min read🧠 Deep dive

Original authors: Jiang, H., Gao, F., Liu, P., Wu, Y., Jie, Y., Li, Y., Jiang, Y.

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). ⚕️ This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer

Imagine a massive, bustling city where millions of people live in distinct neighborhoods, each with its own unique culture and language. Now, imagine you are handed a single, blended recording of the city's total noise—a cacophony of voices, traffic, and music mixed into one indistinguishable stream. Your task is to figure out exactly how many people live in each neighborhood just by listening to that single recording. This is the daily challenge faced by scientists studying the human body. Our tissues are not uniform blocks of cells; they are complex mosaics of different cell types, from immune defenders to structural support workers. To understand how these tissues function in health or disease, researchers often take a sample of tissue, measure the total genetic activity of all the cells inside, and then try to mathematically separate that mixture back into its individual parts. This process, known as deconvolution, is a workhorse of modern genomics, allowing scientists to estimate the proportions of different cell types without having to physically separate them one by one.

However, there is a fundamental problem with this approach. To solve the puzzle of the mixed recording, you need a reference guide—a library of what each neighborhood sounds like when it is alone. In the real world, these reference guides are often incomplete. Scientists might have a perfect recording of the immune cells but lack a clear recording of a rare, fragile cell type that is difficult to isolate. When a reference is missing a piece, the mathematical solution becomes unstable. For years, the scientific community has tried to patch this hole by inventing clever algorithms that guess the missing pieces or by simply ignoring the gap and hoping the remaining data is enough. The prevailing assumption has been that if you have a good enough computer model, you can still get a reliable answer, even if your reference library is imperfect.

A new study challenges this comforting assumption, revealing that the problem is far deeper than a simple lack of data. The researchers, led by a team at Peking Union Medical College Hospital, discovered that incomplete references do not just make the answer slightly less accurate; they make the answer fundamentally unidentifiable. In other words, the data itself does not contain enough information to determine the true proportions of the cells. Worse still, the study shows that the way scientists process the data matters just as much as the data itself. If two researchers use the exact same incomplete reference and the exact same tissue sample, but they have learned their mathematical tools from slightly different versions of that reference in the past, they will arrive at completely different conclusions. One might find that a specific cell type is increasing in a disease, while the other finds it is decreasing. Both are following the rules, both are using the same raw numbers, yet they are telling opposite stories.

The researchers demonstrated this by running a massive series of experiments across nine different human tissue cohorts, including blood, lung, and brain tissue. They took standard reference guides and deliberately removed one cell type at a time to simulate an incomplete library. Then, they ran the analysis twice for every single sample. In the first run, they kept the mathematical "memory" of the tool exactly as it was when it was trained on the complete, perfect reference. In the second run, they forced the tool to re-learn everything from scratch using only the incomplete, damaged reference. The results were startling. In nearly 30% of the cases, the two methods produced results that were so different they could not be reconciled. More critically, in over 160 instances across the nine cohorts, the two methods flipped the direction of a scientific association. A cell type that appeared to be linked to a disease in a positive way under one history appeared to be linked in a negative way under the other. This means that the history of how a tool was trained can change the scientific conclusion drawn from the data, even when the data and the final reference look identical.

The study goes further to show that this is not just a problem of the computer tool, but a problem of the data itself. Even if you could freeze the computer tool and stop it from changing, the missing cell type creates a mathematical void that cannot be filled. The researchers proved that for a given mixture of cells, there are countless different ways to fill in the missing piece that would all fit the observed data perfectly. Some of these ways would make one cell type look dominant, while others would make a different type look dominant. Because the data cannot distinguish between these possibilities, the true answer is hidden. The study found that simply adding more samples or running the same experiment multiple times does not solve this. If the underlying reference is incomplete, repeating the experiment just gives you the same ambiguous answer over and over again.

To address this, the team proposed a new way of thinking about these experiments. Instead of trying to force a single, precise number out of the data, they suggest scientists should report a range of possible answers and acknowledge the uncertainty. They developed a software tool called fitdrop that helps researchers check how much their results depend on the history of their reference library and how much the missing data is driving the uncertainty. The tool can also calculate exactly how precise a measurement would need to be to finally solve the puzzle. For example, in the case of blood cells, the study calculated that to confidently determine the direction of a specific association, the measurement of RNA yields would need to be accurate to within roughly 2.6%, a level of precision that current methods do not always achieve.

The findings serve as a stark reminder that in science, having a lot of data does not always mean you have the answer. If the reference guide is missing a chapter, no amount of computational power can write that chapter for you. The study concludes that the field must move away from treating these estimates as simple, fixed numbers. Instead, researchers must treat them as conditional results that depend heavily on the choices made during the analysis. By being transparent about the history of their tools and the limits of their references, scientists can avoid drawing false certainties from incomplete information. The path forward is not to build bigger, more complex models, but to recognize the boundaries of what can be known and to design experiments that specifically target the missing pieces of the puzzle.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →