How accurate are Bayes factor-based null hypothesis tests? A simulation study
This simulation study validates the accuracy of Bayes factor-based null hypothesis tests in common psychological designs using the brms and bridgesampling packages, demonstrating that estimates are reliable when no algorithmic warnings are issued but should be disregarded when such warnings appear.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are a detective trying to solve a mystery: does a new clue actually prove your suspect is guilty, or is it just a coincidence? In the world of science, especially in fields like psychology and linguistics, researchers often face this same question. They have data, and they have two competing stories: one where a specific effect exists (the "guilty" story) and one where nothing is happening (the "innocent" story). To decide which story is better, scientists use a mathematical tool called a Bayes factor. Think of a Bayes factor as a super-precise scale that weighs the evidence. If the scale tips heavily toward the "guilty" story, the researcher feels confident; if it stays balanced, they know the evidence is weak.
However, calculating the exact weight on this scale is often like trying to count every single grain of sand on a beach while standing on a moving boat. It's mathematically impossible to do perfectly, so scientists use clever computer shortcuts to get an approximate answer. The big worry is: how close is that approximation to the truth? If the computer's guess is off, the detective might convict an innocent person or let a guilty one go free. This is the exact problem tackled in a new study by Daniel J. Schad and Martin Modrák. They wanted to know if the popular computer tools scientists use to weigh these Bayes factors are actually accurate, or if they are secretly broken in ways that could lead to wrong conclusions.
To test this, the authors didn't just look at real data; they built a "simulation lab." Imagine they created thousands of fake experiments where they knew the answer beforehand. They generated fake data using a specific set of rules, then asked the computer tools to calculate the Bayes factor and see if the tool's conclusion matched the truth they had planted. They focused on three very common types of experimental designs used in psychology and linguistics: a simple 2x2 design, a design mixing subjects and items (like different people reading different words), and a more complex version of that mix.
The results were a tale of two very different outcomes, depending on a tiny detail: a warning message. When the computer software ran its calculations and did not show a warning message, the Bayes factors were spot-on. The scale was perfectly calibrated; the tool's estimate of the evidence matched the truth almost exactly. It was as if the detective's scale was working perfectly, giving reliable verdicts every time.
However, the story changed dramatically when the software did issue a warning. In the simulations where the computer said, "Hey, I'm having trouble estimating this number," the results became unreliable. The Bayes factors were no longer accurate; they were biased and wildly variable. Sometimes the tool would exaggerate the evidence, making a weak effect look strong (a "liberal" bias), and other times it would downplay a real effect, making it look like nothing was happening (a "conservative" bias). The authors found that in these warning cases, the computer was essentially guessing with a lot of noise, and the results could not be trusted.
Interestingly, the authors tried to fix this by telling the computer to work harder—running more calculation steps (increasing from 10,000 to 40,000 iterations). While this helped a tiny bit, it didn't stop the warnings or fully fix the bias. The study suggests that the problem isn't just about doing more math; it's about the shape of the data being too tricky for the current shortcuts to handle smoothly.
So, what does this mean for the future of science? The authors conclude that for the common experimental designs they tested, Bayes factors are a great tool—but only if the computer stays silent. If the software flashes a warning, the researchers should stop and not trust the result. It's a bit like driving a car: if the "Check Engine" light is off, you can drive with confidence. But if that light is blinking, you shouldn't just ignore it and keep speeding; you need to pull over and fix the problem before you trust the speedometer. The study serves as a crucial reminder that in the high-stakes game of scientific discovery, paying attention to the computer's warning signs is just as important as the data itself.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.