← Latest papers
💻 bioinformatics

Matched full-UDG and non-UDG ancient DNA libraries reveal trade-offs in post-mortem damage correction for imputation and kinship inference

By comparing matched full-UDG and non-UDG ancient DNA libraries from medieval Mongolian individuals, this study demonstrates that while various computational damage correction methods (trimming, rescaling, and masking) differentially balance site retention against error reduction, the optimal choice depends on the specific downstream analysis, with corrected data proving essential for accurate kinship inference in non-UDG libraries.

Original authors: Ravdandorj, O., Sampildondov, C., Janchiv, K., Gakuhari, T.

Published 2026-09-13
📖 4 min read☕ Coffee break read

Original authors: Ravdandorj, O., Sampildondov, C., Janchiv, K., Gakuhari, T.

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). ⚕️ This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer

When scientists study the DNA of people who lived thousands of years ago, they are working with a material that has been slowly falling apart since the moment of death. Over centuries, the chemical building blocks of this genetic code change. Specifically, a component called cytosine often transforms into uracil, a different molecule that the body does not naturally use. When researchers copy this ancient DNA to read it, the machinery mistakes the altered cytosine for thymine, creating a false signal that looks like a genetic difference where none actually exists. This damage is not random; it tends to cluster at the very ends of the DNA fragments, like fraying on the edges of an old rope. To get a true picture of who these ancient people were and how they were related, scientists must decide how to handle this damage. They can use enzymes to clean the DNA before reading it, or they can use computer programs to ignore or adjust the damaged parts afterward. The question of which approach works best, especially when mixing different types of data, has long been a matter of debate.

A team of researchers in Japan and Mongolia recently tackled this problem by looking at two individuals buried in a medieval cemetery in eastern Mongolia, dating back to the thirteenth or fourteenth century. These two people, a man and a woman, appeared to be closely related, perhaps a parent and child, based on initial screening. The scientists took DNA extracts from the same bones of each person and split them into two groups. One group was treated with an enzyme called full-UDG, which chemically removes the damaged parts of the DNA before sequencing, effectively cleaning the sample. The other group was left untreated, preserving the natural damage patterns that serve as a fingerprint of ancient material. This created a perfect match: the same genetic source, processed in two different ways, allowing the researchers to see exactly how the cleaning process changed the results. They then applied various computer methods to the untreated data to see if they could fix the damage without throwing away too much information.

The study revealed that there is no single perfect way to fix ancient DNA damage; every method involves a trade-off. When the researchers tried to simply cut off the damaged ends of the DNA strands, they lost a significant amount of usable genetic data. In contrast, methods that adjusted the confidence scores of the damaged bases or masked them out kept almost all of the data intact. One specific approach, which adjusted the quality scores of the bases near the ends, retained 99.3 percent of the covered sites, while a method that simply cut off five bases from each end kept only about 85 percent. Interestingly, the damage in these specific samples was not symmetrical; it was much heavier on one end of the DNA strands than the other. When the team tested cutting only the heavily damaged end rather than both ends equally, they found they could keep even more data while still reducing errors. This suggests that for certain types of ancient samples, a one-sided approach to cleaning might be more efficient than a balanced one.

The researchers then used these cleaned and uncleaned datasets to reconstruct the genetic profiles of the two individuals and determine their relationship. The results showed that the choice of cleaning method could shift the estimated closeness of their family tie. When the damaged, untreated data was used without correction, the computer suggested the two were second-degree relatives, like an aunt and nephew. However, once the data was corrected using the best computer methods, the analysis consistently pointed to a first-degree relationship, such as a parent and child. This shift confirmed that the uncorrected damage was misleading the computer into thinking the two people were less related than they truly were. Further analysis of their mitochondrial DNA, which is passed down only from mothers, and their genetic sex confirmed that the woman was the mother and the man was the son.

Ultimately, the study demonstrates that while no single computer method is perfect for every situation, the choice of how to handle ancient DNA damage matters greatly for the final story. The researchers found that methods which adjust the data without discarding it entirely tend to produce the most consistent results when comparing different types of libraries. They also showed that for samples with uneven damage, tailoring the cleaning process to the specific pattern of decay can preserve more valuable information. By carefully balancing the removal of errors with the retention of data, scientists can more accurately reconstruct the lives and family histories of people from the deep past, turning frayed, damaged fragments into a clear picture of human connection.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →