Z-DNA-associated genomic instability in the human pangenome
By leveraging long-read sequencing and human pangenome assemblies, this study reveals that predicted Z-DNA-forming sequences are associated with elevated mutation densities, particularly for insertions and deletions, across diverse human haplotypes and genomic contexts.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer
Inside every human cell, the genetic code is usually stored as a double helix, a twisted ladder structure known as B-DNA. This right-handed spiral is the standard shape we see in textbooks, but DNA is a flexible molecule that can twist into other forms under specific conditions. One of these alternative shapes is called Z-DNA. Unlike the smooth, right-handed twist of the standard form, Z-DNA is a left-handed helix with a jagged, zig-zag backbone. This unusual shape tends to form in specific stretches of genetic code, particularly where two types of building blocks, called purines and pyrimidines, alternate in a repeating pattern. While scientists have long known that Z-DNA plays a role in turning genes on and off and in the immune system, a lingering question remained: does the tendency of DNA to twist into this left-handed shape also make the genetic code more prone to breaking or changing?
For years, answering this question was difficult because the maps scientists used to study human DNA were incomplete. These older maps, like the reference genome GRCh38, had large gaps, especially in repetitive regions where the genetic code repeats itself over and over. These gaps hid the very places where Z-DNA is most likely to form. However, recent advances in sequencing technology have allowed researchers to build complete, end-to-end maps of human chromosomes, including these previously hidden repetitive areas. By using these new, complete maps alongside a massive collection of genetic data from hundreds of diverse people, a team of researchers has now been able to trace the relationship between Z-DNA and genetic instability across the entire human genome with unprecedented clarity.
The researchers began by using a specialized computer program to scan the complete human reference genome and 464 individual human genomes from the Human Pangenome Reference Consortium. This program looked for the specific sequences of genetic code that are most likely to twist into the left-handed Z-DNA shape. They found that while the overall amount of these sequences was similar across different human populations, the new complete maps revealed significantly more Z-DNA-forming regions than the older maps did. This difference was most noticeable in the repetitive sections of the genome, such as the short arms of certain chromosomes and the regions around the centromeres, which were previously missing or incomplete in older references. The study confirmed that these left-handed structures are not randomly scattered; they are concentrated in areas where genes are active and in regions rich in ribosomal RNA, the machinery cells use to build proteins.
Having mapped where these structures exist, the team then asked if these locations were hotspots for genetic errors. They analyzed more than 52 million genetic variations found in the human pangenome data, comparing the frequency of mutations near Z-DNA sequences against carefully matched control regions that looked similar but did not form Z-DNA. The results showed a clear pattern: genetic mutations were significantly more common near the predicted Z-DNA sites. This was true for almost every type of mutation they examined, but the effect was strongest for small and medium-sized insertions and deletions, where pieces of DNA are added or removed. The researchers also looked at complex events where insertions and deletions happen together, finding these were even more likely to occur near the left-handed structures.
The study revealed that this instability is not spread out evenly but is tightly focused. The mutations were most concentrated right at the boundaries where the DNA switches from its standard right-handed shape to the left-handed Z-shape. This suggests that the physical stress of the DNA twisting into this unusual form, or the cellular machinery trying to manage that twist, creates a fragile point where the genetic code is more likely to break or be copied incorrectly. Furthermore, the researchers found that this relationship depended heavily on where the DNA was located in the cell. The link between Z-DNA and mutations was very strong in gene-rich, active areas of the genome but was much weaker or even absent in the tightly packed, inactive regions known as heterochromatin.
To ensure these findings were not just a result of natural selection filtering out bad mutations over generations, the team also examined a set of more than 250,000 brand-new mutations that had just occurred in children and were not yet present in their parents. They found the same pattern: these fresh mutations were also more likely to appear near Z-DNA sequences. This confirmed that the left-handed shape of the DNA itself creates a local environment that increases the chance of genetic change, independent of whether those changes have survived in the population.
The researchers also explored how the strength of the Z-DNA signal influenced these errors. They found that for single-letter changes in the code, the risk of mutation was actually higher in regions with a lower predicted tendency to form Z-DNA, whereas for larger, more complex errors, the risk was highest in regions with a moderate tendency. This suggests that the relationship between DNA shape and genetic stability is nuanced and depends on both the specific type of error and the local genomic environment. The study concludes that while Z-DNA is a known regulator of gene activity, its physical presence also acts as a significant source of genetic variation, particularly in the repetitive and gene-rich parts of our genome that were only recently brought into full view by modern sequencing technology.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.