← Latest papers
💻 computer science

Automated reproducibility assessments in the social and behavioral sciences using large language models

This study demonstrates that large language models can serve as a scalable screening tool for automated reproducibility assessments in social and behavioral sciences, achieving qualitative conclusions comparable to human reanalysts while supporting, rather than replacing, expert judgment in systematic audits.

Original authors: Stefan Feuerriegel, Tobias Holtdirk, Pietro Marcolongo, Anna Steinberg Schulten, Felix Henninger, Stefan Rose, Sarah Ball, Bolei Ma, Frauke Kreuter, Markus Weinmann

Published 2026-08-05
📖 5 min read🧠 Deep dive

Original authors: Stefan Feuerriegel, Tobias Holtdirk, Pietro Marcolongo, Anna Steinberg Schulten, Felix Henninger, Stefan Rose, Sarah Ball, Bolei Ma, Frauke Kreuter, Markus Weinmann

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine the world of science as a giant, bustling library where researchers write books about how people think, act, and interact. For a long time, there's been a nagging worry in this library: sometimes, when other experts try to read the original notes and redo the math from a book, they can't get the same answer. This is called the "reproducibility crisis." It's like if you followed a recipe for chocolate cake, but when you tried it yourself, you ended up with a bowl of soup. In the social and behavioral sciences—which study things like why we vote, how we learn, or what makes us happy—checking these recipes is incredibly hard. It usually requires hiring a team of expert chefs (scientists) to spend weeks or months in the kitchen, trying to figure out exactly which ingredients were used and how they were mixed. It's slow, expensive, and difficult to do for every single book on the shelf.

Enter the new kid on the block: the Large Language Model (LLM). You might know these as the super-smart AI chatbots that can write stories, answer questions, and even write computer code. Think of an LLM as a tireless, hyper-quick apprentice chef who can read a recipe book, look at the ingredients, and instantly try to bake the cake themselves. The big question scientists have been asking is: Can this digital apprentice actually do the job? Can it look at a published study, find the data, run the numbers, and tell us if the original result was real or just a fluke? If it can, it could change the library forever, turning a process that takes years into one that takes minutes, helping us trust the books on our shelves much more.

This paper is the story of a massive experiment to find out if that digital apprentice is ready for the kitchen. The researchers, led by a team from universities in Germany and the US, set up a giant test involving 180 real studies from psychology, economics, and political science. They gave an AI model (specifically a version called Claude Opus 4.7) the original data and the "recipe" (the study's claim) and asked it to recreate the results. They ran this test five times for each study to see how consistent the AI was, just like baking the same cake five times to see if it turns out the same way every time.

The results were a mix of "wow" and "whoa." The AI was surprisingly good at understanding the big picture. In 80% of the cases, the AI agreed with the original study on whether the main idea was true or false. It was like the apprentice saying, "Yep, this cake is definitely chocolate," just like the original chef. However, when it came to the exact measurements—the specific size of the cake or the precise amount of sugar—the AI struggled a bit more. Only about 24% of the time did the AI get the exact "effect size" (a specific number measuring the strength of a result) within a very tight margin of error. In the other cases, the AI's cake was still chocolate, but maybe a little too sweet or a bit too dry compared to the original.

To see how the AI stacked up against real humans, the researchers looked at a smaller group of studies where human experts had also tried to redo the math. Here, the AI actually did quite well, matching the original findings in 40% of cases, which was actually a bit better than the human experts, who matched them in only 28% of cases. The humans and the AI both seemed to be making different choices about how to mix the ingredients, which is normal in science, but the AI was surprisingly good at sticking close to the original recipe.

The team also tested if the AI needed the whole recipe book or just a short summary. They gave the AI the full paper, the paper without the "how-to" section, and just the abstract (the short summary at the start). Surprisingly, it didn't matter much; the AI performed about the same in all three situations. This suggests the AI is really good at guessing the right steps based on the main goal, even if it doesn't have every single detail. They even tried different AI models, including some that are free and open for anyone to use, and they all gave similar results.

So, what's the verdict? The paper suggests that these AI models are not yet perfect master chefs who can replace human experts entirely. Sometimes the AI gets stuck, can't find the right ingredients, or makes a calculation error. But, it is an incredibly powerful tool that can act as a "first-pass" screener. Imagine a librarian who can quickly scan 100 books a day to flag the ones that look suspicious, so the human experts only have to spend their time on the tricky ones. The authors argue that while AI shouldn't be the final judge, it can be a scalable, fast, and helpful assistant to make science more rigorous and trustworthy, helping us catch errors before they become part of our permanent history.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →