Pangenome discovery and characterization of human protein-coding duplicated genes
By integrating long-read genomic and transcriptomic data from diverse human samples, this study characterizes the human pangenome's protein-coding duplicated genes to discover thousands of novel copy number polymorphic genes, refine existing gene models by reclassifying pseudogenes as functional, and reveal that evolutionary constraints are predominantly found in ancestral rather than recently derived duplicated genes.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer
Imagine your body is a massive, bustling library containing the instruction manuals for building and running a human being. For decades, scientists thought they had cataloged every single book in this library, but there was a tricky section: the "Copy-Paste Zone." In this part of the library, entire chapters of instructions were duplicated over and over again, creating stacks of nearly identical pages. Because these pages looked so much alike, the librarians (scientists) couldn't tell which copy was which, so they often just ignored the duplicates or assumed they were broken, useless drafts. This matters because those messy, duplicated sections are actually where the most exciting new stories are being written—places where our bodies might be inventing brand-new tools to help us think, grow, or adapt. To solve this puzzle, scientists needed a better way to read the library, moving from blurry photocopies to high-definition, 3D scans of the books to see exactly how the duplicates differ.
This paper is like a team of super-sleuths finally cracking the code of that messy "Copy-Paste Zone" in human DNA. The researchers used a powerful new toolkit: 298 long-read assembled human genomes (which act like high-resolution maps of our genetic library) and a massive collection of 5.6 billion full-length cDNA transcripts from 83 different tissues (which are like recordings of the library's books actually being read out loud). By combining these, they were able to sort through 493 families of gene duplicates and found something surprising: 2,713 potential genes that are missing from our standard reference map entirely. These aren't just random glitches; they are likely copy-number polymorphic, meaning different people have different numbers of these copies, just like how some families might have extra copies of a favorite recipe while others don't.
When the team looked closely at the gene families where they could tell the copies apart, they found that 60.0% of them are actually active and working, keeping their "open reading frames" (the parts of the code that make functional proteins) intact. Even more interesting, 45.7% of these working genes are most active in the brain, embryos, or testis, suggesting they play a crucial role in our development and complex functions. The team also fixed 386 existing gene models, including 150 that were completely missing or different from the current "gold standard" T2T-CHM13 annotation. They also reclassified 236 genes (35.1%) that were previously thought to be broken "pseudogenes" as actual protein-coding genes, because they found evidence that these genes are being transcribed, have working code, and have accessible promoters (the "on" switches) in the chromatin.
The study also measured how much these genes are "constrained," or how strictly nature protects them from changing. They found that 24.2% of these segmental duplication (SD) genes show signs of being protected from both copy number changes and amino acid mutations, meaning they are likely important for survival. However, there is a twist in the story: the majority of these important, protected genes are ancient, inherited from our distant ancestors. In contrast, only 16.2% of the newer, derived duplicated genes that emerged recently in the human lineage show this same evidence of constraint. This suggests that while our genetic library is constantly adding new, duplicated chapters, most of these fresh additions are still being tested by evolution, and only a small fraction have proven to be essential enough to be strictly guarded. Ultimately, this pangenome approach gives scientists the precision needed to tell the difference between functional genes and broken pseudogenes, highlighting the specific gene innovations that have arisen most recently in human evolution.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.