Donor-allocation effects vary with reference-cell budget in MuSiC PBMC simulations
This study demonstrates that while the accuracy penalty of unequal donor allocation in MuSiC PBMC deconvolution decreases as the reference-cell budget increases, significant variability in individual configurations and persistent discrepancies with flow cytometry in real samples indicate that the trend's transferability to measured bulk data remains unestablished.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer
To understand the human body at a microscopic level, scientists often look at a sample of blood and try to count how many different kinds of cells are present. A single drop of blood contains a complex mixture of immune cells, such as B cells, T cells, and monocytes, each playing a unique role in defending the body. When researchers analyze the genetic material from a whole drop of blood, they get a blended signal, like hearing a choir sing a single chord rather than hearing individual voices. To figure out which singers are present and how many of each, they use a mathematical method called deconvolution. This process relies on a reference library: a collection of genetic profiles from individual cells, ideally gathered from many different people, to teach the computer what each cell type sounds like on its own.
The quality of this reference library matters immensely. If the library is built from cells taken from just a few people, or if those people are represented unevenly, the final count of cells in the blood sample might be wrong. A critical question arises when building these libraries: does it matter if one person contributes many cells to the reference while another contributes very few, as long as the total number of cells remains the same? This question sits at the heart of a recent study by Yunhao Jiang, which investigated how the distribution of cells from different donors affects the accuracy of these cell counts. The research suggests that while having more total cells generally helps, the way those cells are shared among donors still leaves a mark on the results, even when the total number is quite large.
The study focused on a specific tool called MuSiC, a popular method used to estimate cell proportions from bulk genetic data. The researchers wanted to test a specific scenario: if they kept the total number of reference cells fixed, would it hurt the accuracy if they took most of those cells from one donor and very few from the others? To find out, they created thousands of synthetic mixtures using data from peripheral blood mononuclear cells, a common type of immune cell found in blood. They simulated a situation where three donors provided the reference cells for six different cell types. In some simulations, the three donors contributed equally, like a fair split of twenty cells each. In others, the split was heavily skewed, with one donor providing fifty cells while the other two provided only five each. They ran these tests at different total budgets, starting with a small group of sixty cells per cell type and increasing the number up to three hundred.
The results revealed a clear trend that depended on the size of the reference group. When the total number of reference cells was small, skewing the contribution toward one donor made the estimates significantly less accurate. The error increased noticeably when one donor dominated the reference pool. However, as the researchers increased the total number of cells in the reference library, this penalty began to fade. When the budget was raised to three hundred cells per type, the average difference in accuracy between the fair split and the skewed split became very small, almost disappearing. This suggests that having a larger pool of data can help smooth out the problems caused by an uneven distribution of donors.
Yet, the story is not as simple as "more data fixes everything." Even at the largest budget tested, where the average error was nearly zero, the researchers found that the skewed allocation still caused problems in many specific cases. In nearly half of the individual test configurations, the uneven distribution still produced higher errors than the balanced one. The average result looked good only because some cases improved while others got worse, canceling each other out. This means that for any single experiment, relying on an uneven donor mix could still lead to inaccurate results, even if the overall trend suggests it is safe. The researchers also tested whether simply throwing away extra cells from a dominant donor to make the groups equal would help. They found that discarding cells to force a balance actually made the results worse, because it reduced the total amount of information available to the computer.
To see if these findings held up in the real world, the team applied the same methods to actual blood samples from healthy adults, comparing the computer's estimates against physical counts taken by a machine that sorts cells one by one. In this real-world test, the method struggled to identify major cell types correctly, regardless of how the donors were allocated or how many cells were used. The computer failed to detect B cells and T cells in almost every sample, suggesting that the issues in real blood samples are much deeper than just how donors are grouped. The study concludes that while increasing the total number of reference cells reduces the average penalty of uneven donor allocation, it does not eliminate the risk for individual cases. Furthermore, the problems seen in simulations did not translate to a solution for the real-world failures observed in actual blood samples. The work highlights that when building reference libraries, scientists must report not just the total number of cells, but exactly how many came from each person, because the balance of contributions still matters.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.