← Latest papers
💬 NLP

Unleashing Scientific Reasoning for Bio-experimental Protocol Generation via Structured Component-based Reward Mechanism

This paper introduces Thoth, a model trained on the new SciRecipe dataset using a "Sketch-and-Fill" paradigm and a structured component-based reward mechanism to generate precise, logically ordered, and executable biological experimental protocols that outperform existing large language models.

Original authors: Haoran Sun, Yankai Jiang, Zhenyu Tang, Yaning Pan, Shuang Gu, Zekai Lin, Lilong Wang, Wenjie Lou, Lei Liu, Lei Bai, Xiaosong Wang

Published 2026-01-28
📖 4 min read☕ Coffee break read

Original authors: Haoran Sun, Yankai Jiang, Zhenyu Tang, Yaning Pan, Shuang Gu, Zekai Lin, Lilong Wang, Wenjie Lou, Lei Liu, Lei Bai, Xiaosong Wang

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to teach a very smart, but slightly chaotic, robot how to bake a complex cake. You give it a recipe, but instead of giving you a clear, step-by-step list like "mix flour, then add eggs," the robot might say, "You need to bake something delicious, maybe use some flour, oh and by the way, don't forget the oven is hot," or it might list the steps in the wrong order, like telling you to frost the cake before you've even baked it.

This is exactly the problem scientists face when they try to use current AI models to generate experimental protocols (the detailed instructions for lab experiments). The AI often sounds confident but produces instructions that are incomplete, out of order, or scientifically impossible to follow.

This paper introduces a new solution called Thoth, a specialized AI designed to be a reliable "lab assistant." Here is how it works, broken down into simple concepts:

1. The Problem: The "Chatty" Robot

Current AI models are great at writing essays or answering trivia, but they struggle with precision. If you ask them to generate a protocol for a biology experiment, they might hallucinate (make up) ingredients, mix up the order of steps, or give vague instructions. In a real lab, this is dangerous and wastes time. It's like a GPS that tells you to turn left into a wall because it's trying to be "creative" with the route.

2. The Solution: A New "Recipe Book" (SciRecipe)

To fix this, the researchers first built a massive library of real, high-quality scientific recipes. They call this SciRecipe.

  • The Analogy: Imagine taking 12,000 real-world cooking recipes from the world's best chefs, cleaning them up, and organizing them into a perfect database.
  • What it does: This dataset covers 27 different areas of biology (like cell biology, cancer research, etc.). It teaches the AI not just what to say, but how real scientists actually think and work.

3. The Thinking Process: "Sketch-and-Fill"

Instead of letting the AI just "chat" out a response, they forced it to use a specific thinking method called "Sketch-and-Fill."

  • The Analogy: Think of an architect building a house.
    • The Sketch: First, the architect draws a rough blueprint with just the bones of the house (walls, doors, windows) in a strict, structured format. This is the "Sketch" phase. The AI breaks the experiment down into tiny, logical building blocks (e.g., "Action: Mix," "Object: Chemical A," "Amount: 5ml").
    • The Fill: Only after the blueprint is perfect does the architect write the beautiful, descriptive text for the homeowner. This is the "Fill" phase, where the AI turns those strict blocks into natural, readable sentences.
  • Why it helps: This stops the AI from rambling. It ensures the logic is sound before it tries to sound polite.

4. The Strict Judge: The "SCORE" Mechanism

Usually, AI is graded on how well its words match a human's words (like a spelling test). But in science, matching words isn't enough; the actions must be right. The researchers created a new grading system called SCORE.

  • The Analogy: Imagine a strict food critic who doesn't just taste the food; they check if the chef followed the recipe exactly.
    • Step Count: Did the chef use the right number of steps? (Not too many, not too few).
    • Order: Did they crack the eggs before whisking them? If the order is wrong, the cake fails.
    • Fidelity: Did they use "sugar" when the recipe said "salt"?
  • The Result: The AI gets a score based on whether the experiment would actually work in a lab, not just whether it sounds good.

5. The Training: From Student to Master

The AI, named Thoth, was trained in three stages, similar to how a human learns a trade:

  1. Reading: It read thousands of protocols to learn the language of science.
  2. Apprenticeship: It practiced following the "Sketch-and-Fill" method on simple tasks.
  3. Mastery: It used the "SCORE" judge to correct its own mistakes, learning to prioritize safety and logic over fancy language.

The Outcome

When tested, Thoth outperformed even the most famous, expensive AI models (like GPT-5 or Claude).

  • The Result: It generated protocols that were concise, logically ordered, and actually executable in a real lab.
  • The Claim: The paper states that Thoth bridges the gap between "knowing the science" and "doing the experiment," creating a tool that scientists can trust to generate reliable, step-by-step instructions without the hallucinations or errors common in other AI.

In short, the paper claims to have built a "smart lab assistant" that thinks like a scientist, checks its own work like a strict supervisor, and produces instructions you can actually follow in a laboratory.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →