← Latest papers
📄 genetic and genomic medicine

Clinical validation of large-scale functional assays: insights from 2,120 gene-truthset-assay evaluations

This study demonstrates that the clinical evidence points allocatable for functional assay validation vary significantly based on the composition and stringency of the 'truthset' used, highlighting that augmenting ClinVar-based sets with systematically generated 'proxy-clinical' benign missense variants can substantially improve evidence strength and underscoring the urgent need for standardized guidance to ensure consistency in clinical variant classification.

Original authors: Allen, S., Rowlands, C. F., Kuzbari, Z., Garrett, A., Durkie, M., Burghel, G. J., Robinson, R., Callaway, A., Field, J., Frugtniet, B., Palmer-Smith, S., Grant, J., Pagan, J., Johnston, E., McDevitt
Published 2026-07-14
📖 5 min read🧠 Deep dive

Original authors: Allen, S., Rowlands, C. F., Kuzbari, Z., Garrett, A., Durkie, M., Burghel, G. J., Robinson, R., Callaway, A., Field, J., Frugtniet, B., Palmer-Smith, S., Grant, J., Pagan, J., Johnston, E., McDevitt, T., Hughes, L., Yarram-Smith, L., Logan, P., Reed, L., Snape, K., McVeigh, T., Hanson, H., Roth, F. P., Starita, L. M., Fowler, D. M., Villani, R., Spurdle, A. B., Adams, D. J., Findlay, G., Turnbull, C., Cancer Variant Interpretation Group UK (CanVIG-UK),

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). ⚕️ This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer

Imagine you are a judge trying to decide if a new genetic variant is a "villain" (pathogenic) or a "hero" (benign) in the human body. To make this decision, you have a special tool: a functional assay. Think of this assay as a high-tech stress test for a protein, measuring how well it works under pressure. If the protein fails the test, it's likely a villain; if it passes, it's likely a hero.

But here's the catch: before you can trust this stress test to judge new suspects, you have to prove it works on known suspects. You need a "Truthset"—a group of genetic variants that everyone already agrees are definitely villains or definitely heroes. You run the stress test on this Truthset to see if the machine correctly identifies them. If it does, you get "Evidence Points" (EPs) that you can use to confidently classify future unknowns.

This paper is like a massive, chaotic experiment where the authors tried 2,120 different ways to build these Truthsets to see how much the "Evidence Points" change. They looked at four famous cancer genes: VHL, BRCA1, BRCA2, and RAD51C.

The Great Truthset Shuffle

The authors discovered that the rules for building a Truthset are currently a mess. Different expert groups have different ideas on what counts as a "known" villain or hero. Some say, "Only use the most famous, double-confirmed villains!" (high stringency). Others say, "Any villain with a single witness is good enough!" (low stringency).

To see how this chaos affects the results, the authors played a game of "What If?" They built Truthsets using:

  • Different Variant Types: Just the "misspelled" letters (missense variants), or just the "broken" genes (truncating variants), or a mix of everything.
  • Different Confidence Levels: Using only the most famous classifications (3 stars) or including the fuzzy ones (1 star or "soft conflicts" where experts disagree).
  • Different Phenotypes: For the VHL gene, they asked, "Does the villain cause Type 1 cancer, Type 2 cancer, or just any cancer?"
  • The "Proxy" Trick: Since there aren't enough known "benign" missense variants in the public database (ClinVar), they created a massive list of "Proxy-Clinical" benign variants. These are variants they systematically classified as benign using strict computer rules, essentially filling the empty seats at the table with carefully vetted guests.

The Shocking Results

The results were wild. Depending on which Truthset you picked, the Evidence Points you could award for a single assay swung wildly.

  • The Range: For pathogenicity (villainy), the points ranged from 0.0 (no evidence at all) to 5.6 (strong evidence). For benignity (heroism), it ranged from -1.5 to -6.4.
  • The "All-Variants" Trap: When they used Truthsets made of only missense variants (the ones we actually care about classifying), the Evidence Points were usually lower than when they used a mix of broken genes (PTVs) and harmless synonyms.
    • Why? It wasn't because the machine was bad at spotting missense villains. It was simply because there were so few known missense villains and heroes in the database to test against. It's like trying to test a metal detector with only 5 coins instead of 500; your confidence score drops because your sample size is tiny.
  • The "Proxy" Power-Up: When they added those systematically generated "Proxy-Clinical" benign variants to the Truthset, the Evidence Points for pathogenicity improved. Even though the machine made a few more mistakes (lower concordance), the sheer number of test cases (power) was so high that it outweighed the errors. It's like having a huge army of test subjects; even if a few are mislabeled, the massive data volume gives you a stronger overall signal.

What the Paper Says "No" To

The authors explicitly argue against the idea that using a small, ultra-pure Truthset is always better.

  • They show that being too strict (only using 3-star, perfectly agreed-upon variants) shrinks your Truthset so much that you lose Evidence Points.
  • They also warn against using a Truthset that mixes different types of variants (like PTVs and missense) if you are trying to judge missense variants specifically. If the machine is great at spotting broken genes (PTVs) but bad at spotting misspelled ones (missense), mixing them together might hide the machine's weakness. The "All-Variants" Truthset can dilute the specific performance you need to know about.
  • They also highlight a circular trap: If you use a Truthset built from data that already included the results of the assay you are testing, you are cheating. The paper notes that using older ClinVar data (before the assay was published) gives a more honest picture than using the latest data, which might have been influenced by the assay itself.

The Bottom Line

The paper doesn't claim to have solved the problem. Instead, it suggests that the current lack of clear rules is a major hurdle. The authors propose that we can get more Evidence Points (better confidence) by:

  1. Augmenting our Truthsets with systematically generated "proxy" benign variants.
  2. Relaxing the stringency of our Truthset rules just enough to get a bigger sample size, without losing too much accuracy.

They conclude that we urgently need explicit, prescriptive guidelines to stop the chaos. Until then, the "Evidence Points" you get for a genetic test depend entirely on which Truthset you happened to pick out of the 2,120 possibilities. It's a reminder that in science, the answer you get often depends on the questions you ask—and the rules you use to build your test group.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →