Mitigating LLM-based p-Hacking by Preregistering for the Next LLM
This paper proposes and empirically validates a protocol to mitigate p-hacking in LLM-based research by preregistering analysis plans and executing confirmatory tests on the first eligible model released after registration, thereby preventing researchers from overfitting to specific model behaviors.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
The Problem: "Cooking the Books" with AI
Imagine a scientist wants to prove that a new diet pill works. Instead of doing a strict, boring experiment, they use a super-smart AI (a Large Language Model, or LLM) to read people's food diaries and decide if they ate less.
The problem is that these AIs are like very flexible clay. If the scientist doesn't like the AI's answer, they can just squish the clay a little bit:
- Change the question slightly ("Ask if they ate less instead of more").
- Tweak the AI's "mood" (a setting called temperature).
- Change how the AI is allowed to answer.
They keep tweaking these settings over and over until the AI finally says, "Yes, the diet pill works!" This is called p-hacking. It's like a student taking a test, erasing answers, and trying again until they get an A. The result looks real, but it's actually just a trick of the question, not a real discovery.
The Solution: The "Locked Box" Protocol
The authors propose a new rule to stop this cheating. They call it "Preregistering for the Next LLM."
Here is how it works, using a Time-Travel Lottery analogy:
- The Practice Round: The researcher plays around with the current AI models to figure out their experiment. They decide exactly what questions to ask and how to analyze the answers.
- The Locked Box (Preregistration): Before they see the results, they write down their entire plan and put it in a time-locked box. They also write down a list of "Eligible Models" (e.g., "Any new AI from Google, OpenAI, or Anthropic released after today").
- The Wait: They wait. They cannot touch the experiment again. They have to wait for a brand new AI model to be released that fits their list.
- The Reveal: As soon as that new AI arrives, the researcher opens the box and runs their exact plan on this new machine.
Why does this stop cheating?
Because the new AI didn't exist when they wrote the plan. The researcher couldn't have "cooked the books" for a machine that wasn't born yet.
Furthermore, the paper found that AI models are like different people. A trick that makes Person A say "Yes" often fails when you ask Person B the same question. If a researcher tricks the old AI, the new AI usually sees right through it and gives a boring, honest answer.
What the Researchers Did (The Proof)
To prove this works, the researchers actually followed their own rules. They set up two fake experiments where they knew the answer was "nothing happened" (a null result):
- Fake AI Reviews: They asked AIs to guess if scientific reviews were written by humans or AI. (They knew all were human, so any AI saying "AI" was a lie).
- Fake Diet Logs: They asked AIs to guess if people ate fewer calories on Day 2 vs. Day 1. (They knew the days were random, so the answer should be 50/50).
They tried 11 different ways to "trick" the AIs on older models.
- The Result: On the old models, the tricks worked often.
- The Test: They preregistered their plan and ran it on the first new model released after they locked the box.
- The Outcome: The tricks failed in about 73% of the cases. The new AI refused to play along with the old tricks.
The "Stress Tests"
The researchers also asked: "What if a cheater tries to outsmart this system?"
- What if they pick the best trick? Even if the researcher picks the single trick that worked best on the old model, it still failed to work on the new model most of the time.
- What if they only use one company's AI? Even if they only bet on OpenAI or only on Google, the new model from that same company usually broke the trick.
- Does it stop real discoveries? No. If there was a real effect (like a diet pill that actually worked), the new AI would still find it. The system stops the fake tricks without blocking the real truth.
The Catch
The only downside is patience. You have to wait for the next AI model to come out. In their study, this wait was about 10 days on average.
The Bottom Line
The paper argues that to trust AI research, scientists shouldn't just run an experiment once. They should write down their plan, lock it away, and run it on a future AI that hasn't been released yet. Because the new AI is a stranger to the old tricks, it acts as a natural "lie detector," filtering out the fake results and leaving only the real ones.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.