'Truthsets' for clinical validation of large-scale functional assays: Practice recommendations from Cancer Variant Interpretation Group UK (CanVIG-UK)
The CanVIG-UK group established nine guiding principles and seven best-practice recommendations for constructing variant "truthsets" to clinically validate large-scale functional assays, specifically addressing the need for consistent, context-appropriate standards to resolve variants of uncertain significance in cancer susceptibility genes.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer
Imagine you are a detective trying to solve a mystery: Is a tiny typo in a person's DNA a harmless glitch or a dangerous clue pointing to cancer? For years, many of these typos (called "Variants of Uncertain Significance" or VUS) have been stuck in limbo because we didn't have enough evidence to say "guilty" or "innocent."
Enter Large-Scale Functional Assays. Think of these as high-tech "simulators" or "video games" where scientists can test thousands of DNA typos at once to see how they break a protein. It's like running a stress test on a thousand different car engines to see which ones sputter and which ones roar.
But here's the catch: How do you know the simulator is actually good at its job? You need to check it against a "Truth Set"—a group of typos where we already know the answer. This is where the CanVIG-UK team (a massive group of UK cancer genetics experts) stepped in to write the rulebook.
The Golden Rule: Don't Mix Your Apples and Oranges
The paper's biggest, most important finding is a strict rule about what goes into your "Truth Set."
The Rule: If you are testing typos that change a single letter (missense variants), your "Truth Set" must only contain those same single-letter typos.
What the paper explicitly argues AGAINST:
The authors strongly reject the idea of mixing different types of DNA errors together to test your simulator. Imagine you are testing a spell-checker designed to find typos in a novel. If you test it by feeding it a mix of:
- Missing whole pages (Protein-Truncating Variants),
- Extra pages that do nothing (Synonymous Variants), and
- Actual typos (Missense Variants),
...and the spell-checker gets a perfect score, you might think it's a genius. But in reality, it might just be good at spotting missing pages! It could still be terrible at finding the actual typos. The paper suggests that using this "mixed bag" approach risks over-estimating how good the assay really is. It's like passing a driving test by driving on a straight, empty road and then claiming you can handle a snowstorm.
The "Truth" is Hard to Find (and How to Fix It)
The team ran a massive simulation, testing 2,120 different ways to build these Truth Sets. They found that finding enough "known innocent" typos (benign missense variants) is a major problem. It's like trying to find a needle in a haystack, but the haystack is empty.
Because of this shortage, the paper suggests a clever workaround:
- The "Proxy" Trick: If you can't find a clinically confirmed "innocent" typo, you can create a "proxy" one. This is a typo that hasn't been seen in a hospital yet, but computer models and population data suggest it's almost certainly harmless. The paper argues that adding these "proxy" innocent variants to your Truth Set suggests it will give you more confidence in your results, rather than cheating.
The "Score" Game
The paper explains that the "score" your assay gets (called Evidence Points) depends entirely on how strict your Truth Set is.
- Strictness vs. Power: If you demand your Truth Set only have "perfect" known answers (very strict), you might end up with too few numbers to make a strong conclusion.
- The Trade-off: The authors suggest that relaxing the rules slightly (allowing "likely" innocent or "likely" guilty variants) is okay if it gives you more data points. They emphasize that using a "looser" Truth Set won't trick the system into giving a fake high score; it just makes the score more reliable because you have more data to back it up.
The "Circular" Trap
The paper warns about a sneaky trap called "circularity." Imagine a detective who writes a report saying a suspect is guilty, and then uses that report to prove the suspect is guilty.
- If a DNA variant was classified as "guilty" because of a specific test, you cannot use that same test to validate itself.
- The paper recommends trying to filter out these circular cases, though they admit this is hard to do perfectly in the real world.
The Bottom Line
The CanVIG-UK team hasn't "solved" the problem of every single DNA mystery, but they have built a baseline for how to play the game fairly.
- They showed (through their 2,120 simulations) that mixing different types of variants risks giving misleading results by over-estimating assay performance.
- They suggest that using only missense variants, potentially boosted by "proxy" benign variants, is the best way to validate these powerful new tests.
- They warn that while these tests are getting better, we must be careful not to over-hype them or use them on the wrong types of DNA errors.
In short: If you want to know if a specific DNA typo is dangerous, you must test it against other specific typos of the same kind. Don't cheat by testing it against whole missing pages or harmless extra pages, or you'll fool yourself into thinking your detector is smarter than it really is.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.