← Latest papers
🧬 genetics

Significance filtering induces denominator selection bias in local genetic correlation: closed-form characterisation, a gate-aware correction, and an applicability diagnostic

This paper demonstrates that the standard practice of filtering local genetic correlation estimates based on univariate heritability significance introduces severe denominator selection bias, and proposes a closed-form, gate-aware correction and diagnostic to accurately recover unbiased estimates.

Original authors: Paquin, D., Jain, R.

Published 2026-10-03
📖 3 min read☕ Coffee break read

Original authors: Paquin, D., Jain, R.

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). ⚕️ This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer

Scientists have long sought to understand how genetics shapes the complex landscape of human behavior and health. To do this, they often look at two different traits, such as schizophrenia and bipolar disorder, to see if the same genetic variations influence both. When these traits are studied across the entire genome, the result is a single number representing their overall shared genetic basis. However, this broad average can hide important details, much like a national average temperature might miss a freezing winter in one city and a heatwave in another. To find these local patterns, researchers use tools that break the genome into smaller, manageable chunks, estimating how strongly two traits are linked within each specific region. One such tool, widely used in the field, has become the standard for mapping these local connections.

A new analysis by Dana Paquin and Riddhiman Jain reveals that this standard tool contains a hidden flaw that distorts its own results. The tool works by calculating a ratio: it divides the shared genetic influence between two traits by the individual genetic influence of each trait. To ensure stability, the software is programmed to only report a result if both traits show a strong enough signal on their own. The researchers discovered that this safety check, intended to filter out noise, actually acts as a trap. By only keeping results where the individual signals are strong, the software inadvertently selects for cases where random chance has inflated those signals. Because the calculation divides by these inflated numbers, the final result is systematically pulled down, making real genetic connections appear weaker than they truly are. In some cases, the tool discards nearly half of the true signal, leaving researchers with a diluted picture of the biology.

The study goes further to show that this distortion is not random; it follows a predictable mathematical pattern that depends on the strength of the data and how much the two groups of people being studied overlap. When the data is weak or the groups share many of the same individuals, the tool can even manufacture a false connection where none exists, creating a correlation out of shared noise. The researchers developed a way to measure exactly how much the results are being pulled down and created a correction method to fix it. They tested this correction across six different pairs of psychiatric traits and found that it could recover the true strength of the genetic links with high precision, reducing the error by nearly ten times compared to the uncorrected numbers.

However, the researchers also found that this correction cannot be applied blindly to every result. In cases where one of the traits is so weak that its genetic signal is indistinguishable from random noise, the tool's safety filter removes almost all the data, leaving nothing reliable to correct. In these situations, the few results that do appear are not genuine findings but rather artifacts of the filtering process itself. The authors provide a diagnostic test to tell researchers when their data is strong enough to be trusted and when it is too weak to yield meaningful answers. Their work does not suggest that all previous studies are wrong, but it does show that many published numbers are likely underestimates of the true biological reality. By understanding these limitations, scientists can now interpret local genetic maps with greater clarity, knowing exactly where the tool is reliable and where it is merely reflecting the constraints of its own design.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →