← Latest papers
🤖 AI

BenchBench-Protocol: Evaluating Real-World Wet-Lab Protocol Reasoning and Modification

The paper introduces BenchBench-Protocol, a novel benchmark derived from real-world scientist modifications to published protocols that evaluates the ability of large language models to reason about and adapt wet-lab experimental procedures, revealing that current top models achieve only moderate success on these routine but complex tasks.

Original authors: Aditya Sivakumar, Ashu Singhal, Nicholas Larus-Stone, Nithin Parsan

Published 2026-08-26
📖 6 min read🧠 Deep dive

Original authors: Aditya Sivakumar, Ashu Singhal, Nicholas Larus-Stone, Nithin Parsan

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

In a laboratory, a scientist rarely follows a printed recipe exactly as written. The world of wet-lab biology is defined by adaptation. A protocol, which is simply a detailed set of instructions for an experiment, is often a starting point. When a researcher moves that experiment to a new cell type, a different machine, or a slightly different chemical environment, they must modify the steps. This is not just about changing a number; it is about understanding how one change ripples through the rest of the process. If a scientist alters the temperature in the first step, they might need to adjust the time in the third step, or swap a chemical in the final one. Doing this correctly requires a deep, practical understanding of cause and effect, where a small error in the middle can ruin the results at the end. For decades, this kind of reasoning has been a uniquely human skill, relying on experience and intuition that computers have struggled to mimic.

Researchers at Benchling, a company that builds software for scientists, have created a new way to test if artificial intelligence can master this specific type of thinking. They introduced a test called BenchBench-Protocol. Instead of asking a computer to solve a made-up problem or recite facts from a textbook, they gave it a real-world challenge: take a published scientific protocol and modify it for a new situation, just as a human scientist would. The team did not invent these challenges from scratch. They looked at thousands of actual experiments performed by real scientists on the Benchling platform. They found instances where a scientist took a standard, published protocol and changed it to fit their specific needs. By comparing the original instructions with the scientist's modified version, the researchers extracted the exact logic the human used. They then turned these real-life changes into 149 distinct test questions. Each question asks an artificial intelligence to explain how it would adapt a protocol and, crucially, to justify why those changes are necessary.

To ensure these tests were fair and accurate, the researchers did not rely on computers to grade the answers. They brought in human experts. Each of the 149 tasks was reviewed by scientists who hold advanced degrees and have spent at least three years working in a laboratory. These experts checked that the questions made sense, that the rules for a correct answer were scientifically sound, and that the task actually reflected the kind of work a real scientist does. They kept only the tasks that received the highest ratings for quality and relevance. The final set covers nine different areas of biology, from working with proteins and cells to sequencing DNA and growing microbes. The goal was to create a benchmark that measures whether a machine can reason through the downstream consequences of a change, rather than just guessing the right answer.

The researchers then put nine different artificial intelligence models to the test. These models were given the original protocol, a description of the new experimental goal, and access to search tools to find information. They were asked to write a free-form response explaining how to modify the protocol. The answers were then graded against the expert-verified rules. The results showed that while these models have made significant progress, they still have a long way to go. The best-performing model, Claude Opus 5, achieved a score of 59.2 percent. This means it correctly addressed the necessary changes and their consequences in just over half of the cases. The other models scored between 34.1 percent and 47.1 percent. Even the top performer left a large amount of credit unearned, indicating that the test is difficult and that the models are not yet perfect.

The study also looked at what happens when a model is allowed to try the same question multiple times, simulating a scientist who might run a simulation or check a few different options before deciding. When the researchers took the best answer from ten attempts, the scores improved, but the gap between the models changed. Some models that were average on a single try became much stronger when given multiple chances, while others remained consistent. This suggests that different models have different strengths; some are more reliable, while others are more likely to produce a brilliant answer if they get lucky. However, even with ten tries, the best model still did not capture all the available points, proving that the benchmark is not yet "saturated" or too easy for current technology.

The researchers found that the models performed differently depending on the type of task. They were relatively good at tasks that required looking at experimental evidence and drawing a logical conclusion. However, they struggled more with tasks that required "tacit knowledge"—the kind of gut feeling or intuition an expert has about what is the most likely or best approach in a messy, real-world situation. This was particularly true for troubleshooting, where a model needs to figure out why an experiment failed and how to fix it. The models also varied in how much they used their search tools. Some models, like Grok, consumed a massive amount of data by searching the internet extensively, while others were more conservative. The study noted that these high-reasoning settings, which allowed the models to think deeply and search widely, came with a high cost in terms of computing power and time.

Ultimately, this work provides a grounded way to measure progress in artificial intelligence for science. By building the test on real modifications made by real scientists, the researchers avoided the trap of creating artificial problems that might not reflect the complexity of actual lab work. The findings show that while artificial intelligence is becoming a helpful tool for scientists, it is not yet a replacement for human judgment in the lab. The models can recall facts and follow instructions, but they still struggle with the nuanced, step-by-step reasoning required to safely and effectively adapt a scientific protocol. As these tools become more common in research, benchmarks like BenchBench-Protocol will be essential for understanding where they succeed and where they still need to learn. The path forward involves not just making models smarter, but teaching them to understand the chain of cause and effect that defines the scientific method.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →