EthniVar: a source-stratified catalogue of BRCA1/2 germline variants quantifies the classification gap for variants reported in Indian patients
The study introduces EthniVar, a source-stratified catalogue of BRCA1/2 germline variants that quantifies a significant classification gap for Indian-specific variants compared to European and Chinese populations and evaluates a computational consensus framework that, while highly accurate, can only prioritize a minority of uninformative Indian variants for further curation.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer
Every person carries two copies of a set of instructions known as genes, which act as the body's blueprint for building and maintaining itself. Among these, the BRCA1 and BRCA2 genes are particularly important because they help repair damage to DNA, preventing cells from growing out of control and turning into cancer. When these genes carry a harmful change, or variant, the risk of developing breast and ovarian cancer rises significantly. Doctors use genetic testing to find these harmful changes, which can guide decisions about screening, surgery, or medication. However, the ability to interpret these test results depends heavily on the data available to scientists. For decades, the vast majority of genetic research has focused on people of European ancestry. This has created a situation where a genetic change found in a person of Indian ancestry might be labeled as "uncertain" simply because there is not enough data from their specific population to confirm whether it is dangerous or harmless, while the same change in a person of European ancestry might be clearly identified as harmful.
A team of researchers from Plaksha University in India set out to measure exactly how large this gap is and whether modern computer tools can help close it. They created a new, detailed catalogue of genetic variants found specifically in Indian patients with breast cancer, comparing it against similar catalogues for people of European and Chinese ancestry. Their work reveals that the lack of data is not just a minor inconvenience but a substantial barrier to care. They found that for variants reported only in Indian sources, less than forty percent had a clear, definitive classification from global medical databases. In contrast, for variants found only in European or Chinese sources, roughly sixty percent were clearly classified. This means that a patient of Indian ancestry is far more likely to receive a test result that leaves their doctors unsure of the next steps, potentially delaying life-saving interventions or causing unnecessary anxiety.
To address this uncertainty, the researchers turned to a suite of six different computer programs designed to predict the impact of genetic changes. These tools use different methods, such as looking at how the gene has changed over millions of years of evolution or analyzing the 3D shape of the protein the gene creates. By combining the predictions from all six tools into a single consensus score, the team tested whether they could sort through the uncertain variants. When they applied this method to a set of known, clearly classified variants, the system performed with very high accuracy, correctly identifying harmful changes almost all the time. However, when they applied this same system to the 356 uncertain variants found only in Indian patients, the results were more modest. The computer tools were able to provide a confident prediction for only about twenty-one percent of these uncertain cases. While this helped clarify the status of some variants, it left the majority still in a state of uncertainty.
The study also looked at how common these genetic changes are in the general population. Harmful variants that cause serious disease are usually very rare because natural selection tends to remove them from the gene pool over time. The researchers confirmed that the variants their computer tools predicted to be harmful were indeed much rarer in population data than those predicted to be harmless. However, this frequency data did not solve the problem for the most difficult cases: variants that had conflicting reports in medical databases. Even when combining computer predictions with population frequency data, none of the twenty-one variants that had conflicting classifications could be definitively resolved. This suggests that for these specific, tricky variants, computer models and population data alone are not enough to provide a clear answer.
The researchers conclude that while computational tools are useful for narrowing down the list of candidates, they cannot fully replace the need for more diverse human data. The gap in classification is not a failure of the technology but a reflection of the fact that the evidence base itself is incomplete. To truly close this gap, the scientific community needs more genetic studies focused on South Asian populations, including detailed records of family histories and specific functional tests for the variants in question. The team has made their entire dataset and the results of their computer analysis available to the public through an online database, providing a roadmap for future research. Their work highlights that until genetic databases reflect the diversity of the human population, patients from under-represented groups will continue to face a higher rate of uncertain results, leaving them without the clear guidance needed to manage their health risks.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.