← Latest papers
💬 NLP

RegCheck: A tool for structured comparisons between study registrations and papers

This paper introduces RegCheck, a modular, LLM-assisted tool designed to facilitate the structured comparison of study registrations with published papers by combining AI efficiency with human expertise to enhance scientific transparency and reproducibility.

Original authors: Jamie Cummins, Beth Clarke, Ian Hussey, Malte Elson

Published 2026-07-14
📖 6 min read🧠 Deep dive

Original authors: Jamie Cummins, Beth Clarke, Ian Hussey, Malte Elson

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you're a detective trying to solve a mystery. The suspect (a scientific study) has left two clues: a Plan written before the crime happened (the study registration) and a Story written after the crime (the published paper). The big question is: Did the suspect stick to the plan, or did they sneak in secret changes to make the story look better?

For years, checking these two clues has been like trying to find a specific needle in a haystack while wearing thick winter gloves. It's slow, boring, and requires a super-human level of focus. Because it's so hard, most detectives (reviewers and editors) just skip the check, hoping the suspect was honest. But as the authors of this paper point out, that's a risky gamble.

Enter RegCheck, a new digital sidekick designed to help detectives do their job without burning out.

Why Not Just Ask a Robot?

You might think, "Wait, can't we just upload both documents to a fancy chatbot and ask it to compare them?" The authors say no, and here's why:

  1. The "Magic Wand" Problem: If you ask a standard chatbot to compare two texts, the answer changes depending on how you ask the question, what mood the robot is in, or even what it ate for breakfast. One day it might say "Everything matches," and the next day it might say "They are totally different." That's not good for science, where we need the same answer every time.
  2. The Privacy Leak: Uploading secret, unpublished research papers to a public chatbot is like shouting your secret recipe into a crowded square. The company running the bot might keep your data to train their future robots.
  3. The Memory Trap: Chatbots remember what you said earlier in the conversation. If the bot gets confused about the "hypothesis" part, that confusion might spill over and make it think the "results" part is wrong too, even if it's not. The errors pile up like dominoes.
  4. The "Hallucination" Risk: Chatbots love to make things up. If they can't find an answer, they might just invent one and sound very confident about it. You'd have to read the whole paper yourself just to check if the robot was lying, which defeats the whole purpose of using a robot!

The Smart Solution: RegCheck

Instead of just asking a robot to "do it," the authors built RegCheck, which works more like a high-tech assembly line with a human supervisor. It uses a special framework called IDEA (Ingestion, Definition, Extraction, Adjudication) to break the job into tiny, manageable steps:

  • Ingestion (The Librarian): It grabs the Plan and the Story and cleans them up, making sure the text is ready to be read.
  • Definition (The Boss): A human expert (you!) tells the system what to look for. Maybe you care about the sample size, or the specific drug dosage. You are the boss; the robot doesn't decide what's important.
  • Extraction (The Scanner): Instead of reading the whole book, the system uses a "smart highlighter" (called embeddings) to find only the specific sentences in the Plan and the Story that match your boss's instructions. It pulls out the exact quotes, like a detective pinning evidence to a corkboard.
  • Adjudication (The Judge): Only at this final step does the robot (a Large Language Model) step in. It looks at the specific quotes the system found and says, "Do these match?" It gives a verdict: Deviation (they changed something), No Deviation (they stuck to the plan), or Insufficient Evidence (we can't tell because the info is missing).

A Real-World Test Drive

The authors tested this with a fake clinical trial for a diabetes drug. Here's what RegCheck found:

  • The Good: It confirmed that the drug dose, the randomization method, and the start date were exactly as planned.
  • The Bad (The Sneaky Stuff):
    • The Outcome Switch: The plan said they would measure the change in blood sugar levels (a number), but the paper reported the percentage of people who got better (a yes/no). This is a classic trick to make results look stronger, and RegCheck spotted it instantly.
    • The Tightened Rules: The plan said they would exclude people with kidney issues if their numbers were below 30. The paper said they excluded anyone below 45. That's a tiny number change, but it could mean they kicked out sicker patients to make the drug look safer. A human might miss this, but RegCheck flagged it.
    • The Early Exit: The plan said they wanted 300 people, but they stopped at 212. RegCheck caught this "under-enrollment" deviation.

How Sure Are We?

The authors are very confident that RegCheck makes the process faster and easier, but they are honest that it's not perfect yet.

  • They have not proved that RegCheck is 100% accurate.
  • They have not claimed it replaces human experts. In fact, they insist that humans must stay in the loop to make the final call.
  • They suggest that RegCheck is generally very good at finding the "needle in the haystack," but they admit it can make mistakes (like flagging a change that isn't actually a problem, or missing a subtle one).
  • They are currently running tests to see how well it agrees with human experts. Their goal is for RegCheck to agree with humans just as much as humans agree with each other.

The Bottom Line

RegCheck isn't a magic wand that solves science's problems overnight. It's a tool that lowers the barrier to entry. It turns a task that takes hours of boring reading into a quick, structured check that takes minutes.

The authors hope that by making this check easy, authors will do it themselves before submitting papers, and reviewers will actually do it during the review process. It's about making sure that when scientists say, "We did exactly what we planned," they actually mean it. And if they didn't? Well, at least now we have a way to catch them.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →