Critical Values Robust to P-hacking
This paper proposes a model of hypothesis testing that incorporates p-hacking to derive larger, "robust" critical values—calibrated using medical science data to be equivalent to classical values at one-fifth the significance level—thereby ensuring that significant results occur with the desired frequency even when researchers adjust their behavior to new standards.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are a detective trying to solve a mystery. In the world of science, this "mystery" is often a question like, "Does this new medicine actually work?" or "Is this coin fair?" To solve it, scientists use a special tool called a hypothesis test. Think of this test as a very strict judge. The judge starts by assuming the suspect is innocent (this is called the "null hypothesis"). The judge only declares the suspect guilty (rejects the null) if the evidence is so overwhelming that it would happen by pure chance less than 5% of the time. This 5% limit is called the significance level, and the specific number the evidence must beat to be considered "guilty" is the critical value.
But here is the problem: real life isn't always as neat as a courtroom. Scientists, just like anyone else, want to win. They want their findings published, their careers to soar, and their theories to be famous. This creates a huge temptation to "cheat" a little bit, not by lying, but by trying again and again until they get the result they want. This is called p-hacking. It's like rolling a die over and over again until you finally get a six, and then only telling your friends about that one six while hiding all the ones, twos, and threes. If everyone does this, the "innocent" suspects start getting convicted way too often, and the whole justice system of science starts to crumble. This is why the "replication crisis" is happening: many famous scientific discoveries turn out to be false because the original scientists kept rolling the dice until they got lucky.
This paper by Adam McCloskey and Pascal Michaillat tackles this cheating problem head-on. They don't try to stop scientists from rolling the dice; they know that as long as the rewards are high, scientists will keep trying. Instead, they built a mathematical model to figure out how to change the rules of the game so that even if scientists keep rolling, the "innocent" suspects are still protected. They found that the current rules are too easy to beat. To fix this, they propose a new, tougher "critical value" that acts as a much higher bar for the evidence to clear.
Here is how their solution works, using a playful analogy. Imagine a scientist is a gamer trying to beat a level in a video game. The "level" is finding a significant result. In the old system, the game had a low difficulty setting. If the scientist failed the first time, they could just hit "retry" as many times as they wanted until they won. Because they could keep trying, they eventually won almost every time, even if the game was supposed to be impossible to beat by chance. The paper shows that this makes the "win" meaningless.
The authors suggest a new rule: Make the game harder. But not just a little harder. They realized that if you make the game harder, the scientist will just play more times to try to win. So, you have to make the game so hard that even with all the extra tries, the scientist only wins the right amount of times (5% of the time, if that's the goal).
To figure out exactly how hard to make the game, the authors looked at real-world data from medical studies. They found that about 20% of studies are stopped before they are finished because the scientists run out of money, time, or energy. This means there is an 80% chance a scientist can finish a single experiment. Using this number, they calculated that to stop the cheating, the "difficulty" of the game needs to be increased drastically.
Their main finding is a simple, powerful rule of thumb: To fix p-hacking, you need to act as if your significance level is five times smaller than it actually is.
If a scientist wants to claim a result is significant at the standard 5% level, they shouldn't use the usual critical value (which is 1.96 for a two-sided test). Instead, they should use the critical value that corresponds to a 1% level (which is 2.58).
Why does this work? The authors show that when you raise the bar to 2.58, scientists will indeed try more times. They will run more experiments, drop more data points, and tweak more models in their desperate attempt to cross that higher line. But because the line is so high, even with all that extra effort, they will only succeed in crossing it by pure luck about 5% of the time. The "cheating" is still happening, but it no longer breaks the system.
The paper explicitly rules out the idea that we can just tell scientists to "stop cheating" or that we can fix this by simply counting how many times they tried. The authors argue that scientists will always adjust their behavior to whatever rules we set. If we set a rule that says "you can only try three times," scientists will just find a way to try four times. The only way to guarantee the system works is to set a critical value that accounts for the fact that scientists will keep trying until they run out of resources.
The authors are quite sure about their math. They didn't just guess; they built a rigorous model using probability theory and "optimal stopping" (a fancy way of saying "when to quit the game"). They tested their idea with different scenarios, like adding costs to experiments or making later experiments harder to finish, and their solution held up. They also compared their findings to a popular suggestion by other researchers to simply lower the significance level to 0.5%. The authors suggest that while lowering the level is a good idea, their model shows exactly how much to lower it based on how much resources scientists have to keep trying.
In the end, the paper offers a practical shield for science. It suggests that if we want to trust our scientific results again, we need to stop using the old, easy-to-beat numbers. We need to adopt these "robust critical values." For a standard two-sided test, that means moving the goalpost from 1.96 to 2.58. It's a higher bar, yes, and it means fewer "breakthroughs" will be announced every day. But the trade-off is that the breakthroughs we do announce will actually be real, and the "innocent" null hypotheses will finally be safe from being convicted by a lucky roll of the dice.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.