← Latest papers
🧬 biology

Thousands of pathogenic variants in ClinVar show signs of incomplete loss-of-function

This study demonstrates that while specific metrics like LOFTEE, pext scores, and NMD analysis can help distinguish pathogenic from benign variants, thousands of known pathogenic ClinVar variants exhibit signs of incomplete loss-of-function, revealing significant limitations in relying solely on automated filtering rules to exclude such variants.

Original authors: Tatyana E. Lazareva, Yury A. Barbitoff, Yulia A. Smirnova, Andrey S. Glotov

Published 2026-09-04
📖 5 min read🧠 Deep dive

Original authors: Tatyana E. Lazareva, Yury A. Barbitoff, Yulia A. Smirnova, Andrey S. Glotov

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). ⚕️ This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer

In the vast library of human biology, our genes act as instruction manuals for building and maintaining the body. Sometimes, a single typo in these instructions—a change in a letter or a missing word—can cause a disease. Scientists have long believed that the most dangerous typos are those that completely stop the manual from working, effectively silencing a gene. These are known as loss-of-function variants. When a genetic test finds such a "broken" gene, it is usually treated as a smoking gun, the strongest possible evidence that a person has a specific rare disease. This assumption guides how doctors interpret genetic results and how researchers search for the causes of illness. However, biology is rarely as simple as a broken switch. Just as a damaged instruction manual might still be read by a clever editor who skips the broken page or finds a workaround, cells have ways to bypass genetic errors. Sometimes, a gene that looks broken on paper is actually still producing a working protein, or at least enough of one to keep the person healthy. Understanding when a broken gene is truly broken, and when it is merely glitching, is critical for distinguishing between a serious medical diagnosis and a harmless genetic variation.

A team of researchers at the D. O. Ott Research Institute of Obstetrics, Gynaecology, and Reproductology in St. Petersburg set out to test how well current computer tools can spot these "fake" broken genes. They turned to ClinVar, a massive public database that collects genetic variants and labels them as either disease-causing or harmless. The researchers focused on thousands of variants that were predicted to destroy a gene's function. They wanted to see if the tools used to filter out false alarms could successfully separate the truly dangerous variants from the harmless ones. Specifically, they looked for signs that a cell might be rescuing a broken gene, such as skipping over the error to read the rest of the instructions, or using a backup start signal to restart the protein-making process.

The team applied several different tests to the data, checking for things like whether the gene was active in the right tissues, whether the cell would destroy the broken instructions before they could cause harm, and whether the cell could splice the instructions back together correctly. They found that some of these tests were quite good at their job. For instance, a tool called LOFTEE, which assesses the quality of the genetic annotation, did a strong job of flagging harmless variants. It correctly identified that many variants labeled as benign were indeed low-confidence predictions, while the vast majority of confirmed disease-causing variants were high-confidence. Similarly, they found a clear pattern regarding nonsense-mediated decay, a cellular quality control process that destroys faulty genetic messages. Variants that managed to escape this destruction were far more likely to be harmless, while those that were destroyed were more likely to be the cause of disease.

However, the study also revealed a significant complication: thousands of variants that are already labeled as disease-causing in ClinVar showed signs that they might not be fully broken after all. Even among the variants that doctors currently trust as the cause of illness, a substantial number displayed features suggesting the gene might still be working. For example, some of these "pathogenic" variants showed signs that the cell might be able to splice the instructions back together to avoid the error. The researchers noted that while some tools, like those predicting how a gene is spliced, showed promise, they were not perfect. There was a lot of overlap between the scores for disease-causing variants and harmless ones, meaning that a computer score alone cannot definitively say if a variant is dangerous.

One striking example the researchers highlighted involved a specific genetic change in a gene called EYS, which is known to cause a form of blindness. This variant was classified by computer tools as low-confidence and showed low activity scores consistent with the cell's ability to escape quality control mechanisms. Yet, it is a known cause of disease in a significant percentage of patients. This finding underscores a difficult reality: a variant can look like it has a way to escape the damage, but still be powerful enough to cause illness. The study suggests that while the tools used to filter genetic data are helpful, they are not infallible. Relying on them as a simple "yes or no" switch could lead to mistakes, either by dismissing a real disease risk or by worrying about a harmless variation.

The researchers concluded that the field needs to be more careful. They found that while certain metrics, such as the likelihood of a gene being destroyed by the cell's quality control, are useful for sorting variants, others, like the potential for the cell to restart protein production, were less helpful for making broad judgments. The study emphasizes that automated filtering systems, which are often used to quickly sift through thousands of genetic results, may be too blunt. Thousands of variants currently classified as dangerous show signs of being incomplete failures, and many harmless variants show signs of being real failures. This means that geneticists cannot simply trust a computer algorithm to make the final call. Instead, each case requires a nuanced look, understanding that a gene's behavior is complex and that a "broken" label does not always mean the gene is truly silent. The work serves as a reminder that in the intricate machinery of human genetics, what looks like a total breakdown on paper might still be a functioning, albeit imperfect, system in the body.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →