← Latest papers
🌿 ecology

Deep sequencing artificially inflates estimates of microbial diversity

This study demonstrates that deep sequencing artificially inflates estimates of microbial diversity by generating spurious sequence variants, a phenomenon observed in both synthetic no-diversity controls and real microbial amplicons, thereby necessitating caution when comparing samples with vastly different read depths.

Original authors: Henry, L., Laderman, E., Bergelson, J.

Published 2026-09-03
📖 4 min read☕ Coffee break read

Original authors: Henry, L., Laderman, E., Bergelson, J.

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). ⚕️ This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer

To understand the invisible world living on our skin, in our soil, or within our gut, scientists often turn to a technique that acts like a genetic census. They take a sample, isolate the DNA of the tiny organisms present, and use machines to read the genetic code of specific genes that act as unique identifiers for each species. By counting how many different versions of these genes appear in a sample, researchers can estimate how many different types of microbes are living there. This method has revolutionized our view of the microbial world, revealing complex communities that were previously invisible. However, for these counts to be meaningful, the data must reflect reality, not the quirks of the machine doing the counting. If the technology itself creates the illusion of new life where none exists, the entire picture of microbial diversity becomes distorted, leading scientists to believe a community is far more varied than it truly is.

A recent study tackles a troubling pattern where the very act of looking deeper into a sample seems to create more life than is actually there. The researchers began by testing a simple, controlled scenario: they took samples that contained only a single type of genetic sequence, meaning there was zero diversity to begin with. These "no-diversity" samples were created using host genes that do not vary much or using synthetic strands of DNA added to the mix. When these uniform samples were sequenced, the results were startling. Instead of seeing just one type of sequence, the machines reported hundreds of different variants. The problem grew worse as the researchers increased the number of reads, or the total volume of data collected. Once the sequencing went beyond one hundred thousand reads, the number of false variants exploded exponentially. It appeared as though the machine was inventing new species out of thin air, simply because it was looking so hard and so long.

This artificial inflation was not limited to the controlled test tubes. When the researchers applied the same analysis to real microbial samples, such as those containing bacteria or fungi, they found the same pattern. As the number of reads increased, the calculated diversity of the community rose, even when the actual biological makeup of the sample remained unchanged. The study showed that this error affected not just the total count of species, but also the perceived variety within specific groups of organisms. The researchers found that this issue was consistent across different types of genetic markers used to identify microbes, suggesting a fundamental challenge in how current sequencing technologies handle large volumes of data.

The team investigated whether specific adjustments could fix the problem. They tested shortening the length of the genetic sequences read by the machine and tried using a different sequencing platform known as the AVITI Element. While these changes helped reduce the number of false variants, they did not eliminate the problem entirely. The artificial inflation persisted, though in a smaller form. The authors conclude that the most practical solution for now is to be extremely careful when comparing samples that have been sequenced to very different depths. If one sample has been read a million times and another only a thousand times, the first will almost certainly appear more diverse, not because it is, but because the machine has had more opportunities to generate errors. The study suggests that using these controlled, zero-diversity samples as a benchmark can help scientists tune their methods to get closer to the truth, ensuring that the maps of the microbial world they draw are based on what is actually there, rather than what the machine imagines.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →