The Difference Between "Replicable" and "Not replicable" is not Itself Scientifically Replicable
This paper argues that aggregating binary replication verdicts across non-exact experiments fails to reliably distinguish between "replicable" and "not replicable" results because inherent heterogeneity creates irreducible uncertainty and unidentifiable variance, thereby undermining the statistical foundations used to declare a replication crisis.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
The Big Idea: Why We Can't Easily Tell "True" Science from "False" Science
Imagine you are a judge trying to decide if a new recipe is "good" or "bad." You ask 50 different chefs to try making it.
- If 45 chefs say, "It tastes great!" and 5 say, "It's terrible," you might conclude the recipe is replicable (good).
- If 25 say "Great" and 25 say "Terrible," you might conclude it is not replicable (bad).
This paper argues that this way of judging is broken.
The authors, Berna Devezer and Erkan O. Buzbas, claim that in modern science, we cannot reliably tell the difference between a "good" result and a "bad" one just by counting how many times it worked in different labs. Even if you run thousands of experiments, the math says the line between "replicable" and "not replicable" is invisible.
The Core Problem: The "Perfect Copy" Myth
To understand why, we need to look at what a "replication" actually is.
- An Exact Replication: This is like a photocopier. You take a document, and the machine spits out an identical copy. In science, this would mean every single detail (the same people, the same equipment, the same time of day, the same instructions) is identical, and only the random luck of the draw changes.
- The Reality: Exact copies are almost impossible. In real life, replications are like hand-copying a document. One person uses a blue pen, another uses a pencil. One writes on a desk, another on a train. One is tired, one is energetic.
The paper calls this "non-exactness." Because every lab is slightly different, every experiment is slightly different.
The Two Models: The "Shared Secret" vs. The "Independent Gamble"
The authors use two statistical models to explain what happens when we mix these imperfect copies.
1. The Benchmark Model (The "Shared Secret")
Imagine a group of chefs who are all secretly following the same master recipe, but they are all slightly bad at reading it.
- If the recipe is "hard to read" (high non-exactness), even if you get 1,000 chefs to try it, their results will still be messy.
- The Lesson: There is a "floor" to how precise you can be. No matter how many more chefs you add, the uncertainty never goes away because the "messiness" of the recipe itself is the problem, not the number of chefs.
2. The Operational Model (The "Independent Gamble")
This is how science actually works. Each lab runs its own unique experiment and reports a simple "Yes" (it worked) or "No" (it failed).
- The authors show that when you only get a "Yes/No" answer from each lab, you lose all the information about how different the labs were.
- The Analogy: Imagine you are trying to guess the average temperature of a city.
- Scenario A: You ask 100 people to look at the same thermometer. They all see the same number, but maybe they are all looking at it wrong.
- Scenario B: You ask 100 people to look at 100 different thermometers. Some are broken, some are in the sun, some are in the shade.
- If you only ask them "Is it hot? (Yes/No)," you cannot tell the difference between Scenario A and Scenario B. You just get a list of "Yes" and "No." You have no idea if the thermometers are broken or if the weather is just chaotic.
The "Replication Crisis" is a Statistical Illusion
The paper argues that the famous "Replication Crisis" (the idea that half of science is fake) might be a misunderstanding caused by bad math.
- The Trap: Scientists often say, "We tried to replicate 100 studies, and only 36% worked. Therefore, science is in crisis."
- The Reality: The authors say that 36% is just a number that mixes apples and oranges. Because every lab is different (non-exact), the "36%" doesn't tell us if the original science was bad, or if the labs just couldn't agree on how to measure it.
- The Verdict: You cannot look at a list of "Yes/No" results and say, "This result is definitely true" or "This result is definitely false." The data structure simply doesn't contain enough information to make that call.
Why "More Replications" Doesn't Fix It
Usually, we think: "If we aren't sure, let's do more experiments!"
- The Paper's Twist: If the experiments are "non-exact" (different labs, different people, different tools), adding more experiments is like adding more people to a room where everyone is shouting in different languages. You just get more noise.
- The "Effective Sample Size": The authors show that a study with 100 different labs might only have the statistical power of 5 or 6 perfect, identical experiments. The "noise" of the differences between labs cancels out the benefit of having more labs.
The Many Labs 4 Example
The authors tested their theory on a real project called "Many Labs 4," where 17 labs tried to replicate a famous psychology experiment about thinking about death.
- The Result: The labs were all different. Some did it online, some in person. Some used different questions.
- The Math: When they calculated the "messiness" (heterogeneity) of these 17 labs, they found it was so high that the data could not possibly tell them if the original result was true or false. The "Yes/No" answers were too blurry.
- The Conclusion: Even with a massive, well-funded project, the data was too "fuzzy" to declare a winner.
Summary: What Should We Do?
The authors are not saying we should stop doing science or stop replicating.
- Replication is still good for learning how things work, finding the right measurements, and understanding the details.
- Replication is bad when we try to use it as a simple "Pass/Fail" test to declare a whole field of science a "crisis" or a "success."
The Final Metaphor:
Trying to decide if a scientific result is "real" by counting "Yes/No" votes from different labs is like trying to judge the quality of a song by asking 100 people to hum it back to you. Some will hum it in a different key, some will forget the lyrics, and some will hum it faster. If you just count how many people "got it right," you aren't judging the song; you are just judging how well the crowd can mimic the song.
The paper concludes that the tools we are currently using to declare a "crisis" in science are mathematically incapable of doing the job they are assigned. We cannot reliably draw a line between "replicable" and "not replicable" with the current methods.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.