← Latest papers
🤖 AI

Assessing Reproducibility in Evolutionary Computation: A Case Study using Human- and LLM-based Assessment

This paper evaluates reproducibility practices in evolutionary computation over a ten-year period by comparing a manual assessment using a structured checklist against a new LLM-based automated pipeline called RECAP, finding significant gaps in reporting and demonstrating that automated tools can effectively monitor reproducibility.

Original authors: Francesca Da Ros, Tarik Začiragić, Aske Plaat, Thomas Bäck, Niki van Stein

Published 2026-02-10
📖 3 min read☕ Coffee break read

Original authors: Francesca Da Ros, Tarik Začiragić, Aske Plaat, Thomas Bäck, Niki van Stein

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are following a recipe for a legendary chocolate cake you found online. You follow the instructions perfectly, but the cake turns out to be a flat, salty mess. You realize the recipe forgot to mention the oven temperature, didn't say whether to use salted or unsalted butter, and the "secret ingredient" was a link to a website that no longer exists.

This paper is essentially a "Recipe Audit" for the world of high-tech science.

The Problem: The "Secret Sauce" Crisis

In the world of Evolutionary Computation (a type of AI that learns by "evolving" solutions, much like biological evolution), scientists publish papers claiming they’ve found a "perfect recipe" for solving complex problems.

However, the researchers discovered a major problem: many of these scientific "recipes" are incomplete. They might tell you what they achieved, but they forget to tell you the exact "temperature" (settings) or provide the "kitchen tools" (the actual computer code) needed to recreate the result. If you can't recreate the result, the science isn't truly reliable.

The Study: The Great Science Inspection

A team of researchers decided to act like health inspectors for a decade's worth of scientific papers (from 2016 to 2025). They looked at 168 major papers to see how "reproducible" they actually were. They used a Checklist—much like a pilot’s pre-flight checklist—to see if the authors included:

  • The Ingredients: The exact data and parameters used.
  • The Kitchen Setup: What kind of computers they used.
  • The Cooking Instructions: The actual code and step-by-step logic.

What they found:

  • The "Incomplete Recipe" Score: On average, papers only scored about 62% on completeness. It’s like a recipe that tells you how to bake a cake but forgets to tell you how long to leave it in the oven.
  • The Missing Tools: Only about 37% of papers actually provided the "tools" (the code or data) to let someone else try it.
  • Broken Links: Some scientists provided links to their code, but by the time the inspectors arrived, the links were dead—like a recipe pointing to a restaurant that has already gone out of business.

The Innovation: RECAP (The Robot Inspector)

Manually checking thousands of papers is exhausting and slow. To solve this, the researchers built RECAP, an AI-powered "Robot Inspector."

Think of RECAP as a highly trained digital sous-chef. Instead of a human reading every single word, RECAP uses a Large Language Model (like the technology behind ChatGPT) to scan the papers and even "try out" the code in a digital sandbox to see if it actually works.

The researchers found that the Robot Inspector was remarkably good—it agreed with the human experts about 67% of the time. While it’s not perfect, it’s a massive step toward being able to scan thousands of scientific papers instantly to flag the ones that are "unreliable recipes."

The Big Picture: Why does this matter?

Science is built on a foundation of trust. If one scientist claims to have found a cure or a breakthrough, everyone else needs to be able to test it to make sure it’s real.

By creating better checklists and using AI to catch missing information, this paper is helping to ensure that when scientists publish a "recipe" for the future, it actually works when you try to cook it in your own kitchen.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →