← Latest papers
📊 statistics

p-Hacking Inflates Type I Error Rates in the Error Statistical Approach but not in the Formal Inference Approach

This paper argues that while p-hacking inflates Type I error rates within the error statistical approach by comparing actual familywise error rates to nominal rates, it does not do so in the formal inference approach because the actual familywise error rate is irrelevant to inferences about the specific, selectively reported hypotheses.

Original authors: Mark Rubin

Published 2026-03-04
📖 6 min read🧠 Deep dive

Original authors: Mark Rubin

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

The Big Picture: Two Ways to Look at "P-Hacking"

Imagine you are a detective trying to solve a crime. You have a list of 20 suspects. You interrogate all 20, but only one of them happens to look guilty by pure chance (maybe they just had a nervous tic). You decide to write a report saying, "Suspect #13 is guilty!" while ignoring the fact that you interrogated 19 other innocent people first.

In statistics, this is called p-hacking. It's when researchers run many tests, hide the ones that didn't work, and only show the one that looks significant.

The big question is: Does this trickery make the math wrong?

Mark Rubin's paper says the answer depends on which "philosophy" of math you believe in. He compares two different schools of thought:

  1. The Error Statistical Approach (The "Real-World Detective")
  2. The Formal Inference Approach (The "Strict Rulebook Lawyer")

1. The Error Statistical Approach: The "Real-World Detective"

The Philosophy: "We must look at the whole story, not just the ending."

Imagine a detective who knows that if you interrogate 20 people, you are almost guaranteed to find one who looks suspicious just by luck. If that detective only shows you the photo of the one "guilty" person and hides the photos of the 19 innocent ones, the detective is lying about the odds.

  • The Analogy: Think of a fishing net. If you cast a net 20 times, you are likely to catch a fish eventually, even if the water is empty. If you only show you the one fish you caught and pretend you only cast the net once, you are inflating your success rate.
  • The Verdict: In this view, p-hacking IS a problem. It "inflates" the error rate. The math says, "If you cast the net 20 times, your chance of a false alarm is high (about 64%)." But the researcher claims, "I only cast it once, so my chance of a false alarm is low (5%)."
  • The Result: The "Actual" error rate is much higher than the "Reported" error rate. The evidence is considered "bad" because the detective hid the context of the hunt.

2. The Formal Inference Approach: The "Strict Rulebook Lawyer"

The Philosophy: "We only judge what is written on the official document."

Now, imagine a lawyer who only cares about the final contract signed in court. The lawyer doesn't care about the messy negotiations, the drafts, or the 19 other suspects the detective interrogated behind closed doors. The lawyer only looks at the final, signed statement: "Suspect #13 is guilty."

  • The Analogy: Think of a lottery ticket.
    • You buy 100 tickets (tests).
    • 99 are losers.
    • 1 is a winner.
    • You throw away the 99 losers and only show the winner to your friend.
    • Your friend asks, "What are the odds of this ticket winning?"
    • The Lawyer says: "Well, this specific ticket has a 1 in 1,000,000 chance of winning. That math is correct."
    • The Detective (Error Statistician) says: "But you bought 100 tickets! You were guaranteed to find a winner eventually!"
  • The Verdict: In this view, p-hacking does NOT inflate the error rate of the specific test reported.
    • The researcher reported a test for "Suspect #13."
    • The math for "Suspect #13" is still valid (5% chance of error).
    • The fact that they looked at 19 other suspects first is a "psychological" issue, not a mathematical one. The math for the specific claim they made remains untouched.
  • The Result: The "Actual" error rate is irrelevant because the researcher didn't claim to be testing all 20 suspects together. They only claimed to test #13. Therefore, the math holds up.

Why the Disagreement Matters

The paper argues that these two approaches are talking past each other because they define "the test" differently.

  • The Detective (Error Statistician) says: "The test includes the process of hunting for the result. If you hunted for 20 hours, that's part of the test."
  • The Lawyer (Formal Inference) says: "The test is only the final claim you wrote down. If you wrote down a test for Suspect #13, that is the only test that exists in the official record."

The "Texas Sharpshooter" Analogy

The paper uses a famous story to explain the difference: The Texas Sharpshooter.

Imagine a guy shoots a bunch of random holes in a barn wall. Then, he paints a target around a cluster of holes and says, "Look! I'm a master sharpshooter!"

  • The Detective's View: This is fraud. He cheated by painting the target after the shots. The "actual" skill is terrible. The error rate is 100% because he rigged the game.
  • The Lawyer's View: If the sharpshooter says, "I am testing my ability to hit this specific painted target," and we assume the target was there before he shot (hypothetically), then the math of hitting that target is still valid. The fact that he painted it after the shots is a lie about his skill, but it doesn't change the mathematical probability of hitting a target if one existed.

Crucial Point: The Lawyer admits the sharpshooter is a fraud regarding his skill (substantive inference), but the math regarding the specific target hit remains logically sound.

What About the "Replication Crisis"?

Why do scientific studies fail to replicate (work again) so often?

  • The Detective's Explanation: "Because p-hacking inflated the error rates! We have too many false alarms!"
  • The Lawyer's Explanation: "No, the math isn't broken. The problem is that we are misinterpreting what the math means. We think a 'significant' result proves a theory is true, but it might just be a random fluke, or we might be ignoring hidden factors (like the color of jelly beans causing acne only in green, not red). The math is fine; our understanding of the real world is flawed."

Summary for the Everyday Person

  • p-Hacking is like fishing in a lake, catching 19 empty nets, and only showing people the one fish you caught.
  • Error Statisticians say: "You cheated! You fished 20 times, so your success rate is fake."
  • Formal Inferentialists say: "You didn't cheat the math. You just hid the empty nets. The math for that one fish is still correct. The problem is that you are misleading us about how hard you tried, not about the fish itself."

The Takeaway: Rubin suggests that while p-hacking is bad for science because it hides the truth about how we found a result, it doesn't necessarily break the math of the result we found. The solution isn't just better math (like preregistration), but better critical thinking about what our results actually mean and admitting that we might be missing hidden factors in the real world.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →