← Latest papers
📄 evolutionary biology

A Rarefaction Approach to Identify Local Introgression in a Three Population Tree

This paper introduces DSTAR, a new rarefaction-based method that overcomes the limitations of the traditional D statistic by enabling precise detection of local introgression in multi-lineage datasets without requiring an outgroup, as demonstrated by its successful identification of Denisovan immune-related DNA in modern Papuans.

Original authors: Smith, T. Q., Szpiech, Z. A.

Published 2026-09-11
📖 7 min read🧠 Deep dive

Original authors: Smith, T. Q., Szpiech, Z. A.

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). ⚕️ This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer

The story of human evolution is not just a straight line of ancestors leading to us; it is a tangled web where different groups of people met, mixed, and exchanged genetic material. This process, known as introgression, happens when two distinct populations interbreed, leaving behind a trail of shared DNA that can be found in their descendants today. For decades, scientists have used a specific mathematical tool to detect this ancient mixing across entire genomes. This tool works by looking for an imbalance in how genetic variants are shared between three groups of people and a fourth, distant group that serves as a reference point to determine which genetic changes are new and which are old. While this method has been excellent at spotting large-scale mixing events, it struggles when researchers try to find the specific, small chunks of DNA where this mixing occurred, especially when the data is messy or the groups being compared have different numbers of individuals.

In a new study, researchers T. Quinn Smith and Zachary A. Szpiech have developed a fresh approach to solve these problems. They created a new method designed to pinpoint exactly where ancient DNA entered a modern population, even when the data is imperfect or the sample sizes are uneven. By applying a statistical technique that adjusts for the number of individuals in each group, their new tool can identify these genetic footprints with greater accuracy than previous methods. The team tested their approach extensively using computer simulations that mimic the complex history of human populations. They found that their method is more reliable at finding the right spots and less likely to make mistakes, even when the genetic data is incomplete or when the groups being studied have evolved at different speeds. To prove its real-world value, they applied their new tool to the DNA of modern Papuan people and successfully identified specific segments of DNA inherited from Denisovans, an ancient human relative, including genes that play a crucial role in the immune system.

The core challenge the researchers addressed is the difficulty of distinguishing between genetic similarities caused by ancient mixing and those caused by other factors, such as random chance or differences in how populations grew and shrank over time. The traditional method, widely used in the field, relies on counting how often specific genetic patterns appear in a set of four groups. It requires a "reference" group, usually an ancient or distant population, to tell the researchers which genetic variants are the original ones and which are newer mutations. However, this reference group is not always available, and even when it is, the method often fails when applied to small sections of the genome or when the number of people sampled from each group varies significantly. If one group has many more individuals than another, the traditional tool can be tricked into seeing mixing where there is none, or missing it entirely.

Smith and Szpiech's solution involves a concept called rarefaction, a technique borrowed from ecology that is used to compare species diversity when sample sizes differ. In their new method, they do not need a reference group to determine the age of a genetic variant. Instead, they look at the unique genetic signatures that appear in pairs of populations. They calculate the probability of finding a specific genetic variant in a smaller, standardized sample drawn from each group. By doing this, they level the playing field, ensuring that a group with many samples is not unfairly weighted against a group with few. This allows them to count how many genetic variants are shared exclusively between two groups, which serves as a strong signal that those groups exchanged DNA in the past.

The researchers put their new method through rigorous testing using computer simulations that recreated the history of human populations. They created scenarios where populations split apart, mixed at specific times, and then continued to evolve. They tested their method against several existing tools, including the traditional approach and other modern techniques that rely on complex computer models. In these simulations, the new method consistently outperformed the others. It was better at correctly identifying the specific chunks of DNA that had been mixed in, and it made fewer false alarms. Crucially, the new method remained accurate even when the researchers removed the reference group entirely, a step that caused the traditional method to lose its power. It also held up well when the data was "depolarized," meaning the researchers scrambled the information about which genetic variants were old and which were new, a situation that often happens in real-world studies where the true history is unknown.

The study also examined how the method would handle the messy reality of ancient DNA. DNA from old bones is often damaged, with missing pieces or chemical changes that make it look like a different genetic variant than it actually is. The researchers simulated these errors, including missing data and chemical damage, and found that their new method was robust. While the presence of missing data did reduce the ability of all methods to find the truth, the new approach remained more reliable than its competitors. It was particularly effective when the number of individuals sampled from different groups was uneven, a common situation in genetic studies where some populations are well-sampled and others are not.

To demonstrate the practical application of their work, the team turned to real-world data from modern-day Papuan people and a sample of Denisovan DNA. Papuans are known to carry a significant amount of DNA from Denisovans, an ancient human group that lived in Asia and Oceania. The researchers used their new tool to scan the genomes of 17 Papuan individuals and one Denisovan sample, looking for the specific regions where this ancient mixing occurred. They identified several blocks of DNA that were likely inherited from Denisovans. Many of these blocks contained genes related to the immune system, such as RRM1, HIC2, BANK1, CLEC9A, and LRBA. This finding aligns with the idea that mixing with ancient humans provided modern populations with genetic advantages, particularly in fighting off local diseases. They also found a gene called ADK, which has been linked to adaptation to high altitudes and is also involved in immune function.

The researchers also tested how well their method would work if they had fewer samples to work with, a common constraint in genetic research. They repeated the analysis using only two Papuan individuals instead of the full group of 17. While the traditional method and other tools lost many of the identified regions in this smaller dataset, the new method retained a much larger proportion of the findings. This suggests that the technique is particularly valuable for studies where sample sizes are limited or uneven, allowing researchers to extract more information from smaller datasets than was previously possible.

The study concludes that while the new method is a powerful tool for finding local regions of ancient mixing, it is best used as a way to generate candidates for further study. It excels at highlighting the most interesting parts of the genome, but it does not replace the need for a full understanding of a population's history. The authors note that the method works best when researchers are looking for outliers—regions that stand out from the rest of the genome—rather than trying to measure the total amount of mixing across the entire genome. In cases where the evolutionary history of the groups is very complex or involves rapid changes in population size, other methods might still be necessary. However, for the specific task of finding the precise locations where ancient DNA entered a modern population, this new approach offers a significant improvement in accuracy and reliability, opening the door to a clearer understanding of our shared genetic past.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →