← Latest papers
📊 statistics

Sensitivity Analysis for False Discovery Rate Estimation with Published p-Values

This paper provides closed-form expressions to analyze how publication bias models, such as the p<0.05p<0.05 rule, impact the bias and variance of false discovery rate (FDR) estimators, demonstrating through empirical examples that these formulas enable effective sensitivity analysis for FDR estimation when the true selection mechanism is unknown.

Original authors: Tianyu Cao, Sangyoon Yi, Joshua Habiger

Published 2026-03-03
📖 5 min read🧠 Deep dive

Original authors: Tianyu Cao, Sangyoon Yi, Joshua Habiger

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are a detective trying to solve a mystery: How many of the "solved cases" (scientific discoveries) published in the news are actually false alarms?

In the world of science, researchers run experiments to find new truths. They publish their results, but there's a catch: Journals only print the "winners." If a study finds a "statistically significant" result (usually a p-value less than 0.05), it gets published. If it's a dud, it goes into the "file drawer" and is never seen. This is called Publication Bias.

Because of this, the pile of published science looks like a gold mine, but it might actually be full of fool's gold. The False Discovery Rate (FDR) is the percentage of these published "gold" that is actually fake.

The Problem: Guessing the Rules of the Game

For years, scientists have tried to calculate this FDR using only the published papers. But to do the math, they have to guess how the papers got selected.

Most people assume the rule is simple: "If the p-value is under 0.05, it gets published. If it's 0.06, it gets rejected."

But is that true?

  • Maybe a p-value of 0.049 is published 100% of the time, but a p-value of 0.051 is published 50% of the time?
  • Maybe a p-value of 0.001 is published 100% of the time, but a p-value of 0.04 is only published 10% of the time?

If you guess the wrong rule, your calculation of "how many fakes are in the pile" will be wrong. You might think the science is 90% true when it's actually only 50% true, or vice versa.

The Solution: A "What-If" Calculator

This paper by Cao, Yi, and Habiger is like building a universal "What-If" calculator for scientific truth.

Instead of forcing everyone to agree on one single rule for how papers get published, the authors created a set of mathematical formulas (closed-form expressions) that let you plug in different selection rules and see how the answer changes.

Think of it like tuning a radio:

  • The Old Way: You tune to one station (the "p < 0.05" rule) and hope it's the right frequency. If you're slightly off, the music is static.
  • The New Way (This Paper): You have a dial that lets you slide the frequency up and down. You can ask: "What if the selection rule is strict? What if it's loose? What if it favors tiny p-values?" The calculator instantly tells you how the estimated "False Discovery Rate" shifts with every turn of the dial.

The Key Findings (The "Aha!" Moments)

  1. The "Conservative vs. Liberal" Trap:
    Depending on how you guess the selection rule, your estimate of false discoveries can swing wildly.

    • If you guess the rule is too strict, you might think the science is safer than it really is (Conservative).
    • If you guess the rule is too loose, you might think the science is riskier than it really is (Liberal).
    • Analogy: Imagine a fishing net. If you think the net has tiny holes (strict selection), you assume you caught only the biggest, best fish. But if the net actually has huge holes (loose selection), you might have caught a lot of junk too, and your estimate of "quality fish" is wrong.
  2. The "Negative Bias" Surprise:
    Usually, we expect our estimates to be a little too high (overestimating the truth). But the authors found that under certain weird selection rules (where very small p-values are super likely to be published, but p-values near 0.05 are rarely published), the estimate can actually be too low. It can tell you "Everything is great!" when actually, the pile is full of junk.

  3. Real-World Testing:
    The authors tested their calculator on real data from psychology and medical journals.

    • The Psychology Test: They looked at a famous project where scientists tried to replicate old studies. They found that their new formulas could accurately predict the "true" error rate, even when they didn't know the exact selection rules the original authors used.
    • The Medical Test: They looked at medical abstracts. They showed that if you assume the selection rule is slightly different (e.g., maybe p-values between 0.025 and 0.05 are less likely to be published than those under 0.025), the estimated error rate jumps from 18% to nearly 30%.

Why Should You Care?

This paper doesn't just give you a new number; it gives you honesty.

In the past, a scientist might say, "Based on my model, 14% of medical discoveries are false."
Now, with this paper, a scientist can say:

"If the selection rule is strict, the error rate is 14%. But if the selection rule is a bit more complex, the error rate could be 29%. Here is a range of possibilities so you can make a better decision."

The Bottom Line

Science is messy, and the way papers get published is even messier. This paper provides a sensitivity analysis toolkit. It admits that we don't know the exact rules of the game, so instead of guessing one answer, it gives us a map of all the possible answers.

It's like telling a policymaker: "We can't be 100% sure how many fake discoveries are in the pile, but here is the range of possibilities. If you assume the worst-case scenario, the problem is huge. If you assume the best, it's manageable. Now you can decide how much caution to apply."

This helps us stop pretending we know the exact truth and start making decisions based on a realistic understanding of uncertainty.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →