PRS-CARV: A summary statistics framework for integrating annotation-informed rare variants to improve polygenic risk prediction
The paper introduces PRS-CARV, a summary statistics framework that integrates annotation-informed rare variant burden scores with common-variant polygenic risk scores to significantly improve the prediction of lipid traits and advance precision medicine.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer
Imagine trying to predict a person's future health by looking at their genetic blueprint. For years, scientists have focused on the most common letters in that blueprint, the tiny variations in DNA that appear frequently across the entire human population. These common variations act like a gentle, widespread breeze, each contributing a tiny amount to a person's risk for conditions like heart disease or high cholesterol. By adding up these small effects, researchers can create a "polygenic risk score," a single number that estimates an individual's genetic susceptibility. However, this approach has long ignored the rare, powerful gusts of wind: the genetic variations that occur in fewer than one in a hundred people. These rare variants are often much more potent, capable of causing significant biological changes, but because they are so uncommon, they have been difficult to study without massive, expensive databases of individual genetic codes.
A team of researchers has now developed a new method to bring these rare, powerful genetic signals into the prediction mix without needing access to those massive, private databases. Their approach, called PRS-CARV, allows scientists to combine the steady influence of common genetic variations with the sharp impact of rare ones, using only publicly available summary data. By testing this method on real-world data from pregnant women, the team found that adding rare variants significantly improved the ability to predict levels of cholesterol and other fats in the blood. This breakthrough suggests that a more complete picture of genetic risk is possible, one that captures both the common background noise and the rare, high-impact signals hidden in our DNA.
The core challenge the researchers addressed was a logistical one. To accurately measure the effect of rare genetic variants, scientists traditionally needed to gather the raw, individual genetic data from hundreds of thousands of people. This is often impossible due to privacy laws, high costs, and the sheer difficulty of sharing such sensitive information. Meanwhile, large-scale projects like the UK Biobank have already analyzed millions of genetic samples and published the results as summary statistics—aggregated numbers that show how specific genes or variants are linked to traits, without revealing who the individuals are. The new method, PRS-CARV, was designed to bridge this gap. It takes the summary data from these large public resources and combines it with the individual genetic data of a specific group of people to build a more accurate risk score.
The researchers tested their framework by looking at lipid traits, which are measures of fats in the blood such as high-density lipoprotein (HDL), low-density lipoprotein (LDL), triglycerides, and total cholesterol. They used a dataset of nearly 3,000 pregnant women of European ancestry from the nuMoM2b-Heart Health Study as their target group. For the base data, they relied on summary statistics from the UK Biobank, which included information on over 390,000 people. The process involved two main steps. First, they calculated a risk score based on the common genetic variants, a standard practice in the field. Second, they calculated a separate score for the rare variants. To do this, they grouped rare genetic changes within specific genes based on their function, such as whether they were likely to break a protein or just change it slightly. They then summed up the effects of these rare groups using the weights provided by the large public database. Finally, they combined these two scores into a single, unified prediction model.
The results showed a clear advantage to including the rare variants. When the researchers looked at how well the scores predicted actual cholesterol levels in the women, the model that included both common and rare variants performed better than the model using common variants alone. The improvement was not uniform across all traits; it ranged from a 0.02 increase in AUC for HDL cholesterol to an 0.11 increase for LDL cholesterol. In every case, the rare variants added a layer of information that the common variants missed. The researchers also observed that the genes contributing to the rare variant score were largely different from those in the common variant score. While the common variants pointed to a broad network of genes with small effects, the rare variants highlighted specific genes where the genetic changes were more likely to disrupt protein function, such as those involved in breaking down fats.
To ensure their findings were robust, the team also ran computer simulations where they created fake genetic data with known properties. In these simulations, they knew exactly which genetic variants caused the traits and how strong their effects were. The simulations confirmed that while rare variants alone were not strong enough to predict the traits well, they provided a significant boost when added to the common variants. This held true even when the researchers changed the number of genetic causes in the simulation. The consistency between the simulation and the real-world data gave the team confidence that their method was working as intended.
However, the study also revealed important boundaries to where this method works. When the researchers applied the same approach to binary outcomes—conditions that are either present or absent, such as gestational diabetes or high blood pressure during pregnancy—the addition of rare variants did not improve the prediction. The model performed no better with rare variants than without them. The authors suggest this is likely because the large public database they used did not have enough cases of these specific pregnancy-related conditions to detect rare genetic effects with statistical certainty. Additionally, these pregnancy conditions are influenced by many complex factors beyond just genetics, such as hormones and the environment, which may dilute the signal from rare variants.
The study concludes that PRS-CARV offers a practical way to make genetic risk prediction more comprehensive. By leveraging publicly available summary statistics, the method allows researchers to incorporate the powerful information contained in rare genetic variants without needing to access private, individual-level data. This is a significant step toward precision medicine, where understanding a person's full genetic risk profile, including both common and rare elements, can lead to better prevention and care strategies. The researchers note that while their current work focused on European ancestry and protein-coding regions of the genome, future applications could expand to include diverse populations and non-coding genetic regions as more data becomes available. For now, the method stands as a proven tool for refining how we understand the genetic architecture of complex traits like cholesterol levels.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.