Measuring Mid-2025 LLM-Assistance on Novice Performance in Biology
A pre-registered, randomized controlled trial conducted in mid-2025 found that while large language models did not significantly increase the overall completion rate of complex viral reverse genetics workflows by novices compared to internet search, they provided a modest performance benefit in individual task success and progression through intermediate steps, highlighting a critical gap between AI's in silico benchmarks and its real-world utility in physical laboratory settings.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are teaching a group of complete beginners how to build a complex, custom robot from scratch. You have two groups of students:
- The "Google Group": They can use the internet, Wikipedia, YouTube, and forums to find instructions.
- The "Super-Brain Group": They get everything the Google Group has, plus access to the world's most advanced AI assistants (the "Super-Brains" of mid-2025).
The big question was: Does having a Super-Brain AI help the beginners build the robot faster and better than just using the internet?
This paper is the report card from a massive, 8-week experiment where 153 people tried to do exactly this, but instead of robots, they were building viruses (specifically, a process called "reverse genetics," which is like reverse-engineering a virus from a digital code). This is a high-stakes test because if AI makes it too easy for beginners to build dangerous viruses, that's a major safety risk.
Here is the breakdown of what happened, using some simple analogies:
1. The Setup: A "Blind" Cooking Contest
The researchers set up a controlled kitchen (a safe biology lab). They didn't give the students a recipe book. Instead, they just said, "Make a virus."
- The Goal: Complete five specific steps, like growing cells in a petri dish, mixing DNA, and making the virus.
- The Rules: The "Google Group" could search for answers. The "Super-Brain Group" could ask the AI for answers. Neither group had a human teacher to help them if they got stuck.
2. The Main Result: The AI Didn't Win the Race
The Headline: The Super-Brain group did not finish the whole project significantly more often than the Google group.
- The Analogy: Imagine a marathon. Both groups were running through a thick fog. The AI group had a GPS, but the Google group had a really good map. In the end, almost nobody finished the race. Only about 5-6% of people in both groups successfully completed the entire virus-building process.
- Why? Building a virus in a real lab is incredibly hard. It requires "muscle memory" and "feel" (like knowing exactly how much pressure to apply to a pipette or how healthy a cell looks). AI is great at giving text instructions, but it can't hold the pipette for you.
3. The Silver Lining: The AI Helped with the "First Steps"
While the AI didn't help them cross the finish line, it did help them get started and move through the middle parts faster.
- The Cell Culture Win: In one specific task (growing cells), the AI group was about 15% more successful than the Google group.
- The Analogy: Think of the AI as a very patient tutor. When the Google group was stuck trying to figure out "What do I do next?", they spent hours searching YouTube and forums. The AI group asked the AI, "How do I feed these cells?" and got a clear, step-by-step answer immediately.
- The Result: The AI group started tasks sooner, made fewer mistakes in the beginning, and got further down the line. They were like hikers who found the trailhead faster, even if the mountain was still too steep for them to reach the summit.
4. The "Trust Gap": Why the AI Struggled
The study found a funny and important problem: The students didn't know how to talk to the AI effectively.
- The Analogy: Imagine you have a genius chef (the AI) who knows every recipe in the world. But you are a beginner cook who doesn't know what "sauté" means or how to describe a broken pan. You ask the chef, "Fix my soup," and the chef gives you a recipe that requires a tool you don't have.
- The Reality: The students often asked the AI for the wrong ingredients or got confused by the AI's suggestions. Sometimes the AI gave them the wrong DNA sequence. The students didn't have enough knowledge to realize the AI was lying or confused.
- The YouTube Surprise: The students actually found YouTube videos more helpful than the AI. Why? Because biology is visual. Seeing a video of someone doing the task (the "tacit knowledge") was better than reading a text description from an AI.
5. The Big Takeaway: "Smart" vs. "Skilled"
This study is a reality check for the future of AI safety.
- The Myth: "AI is so smart it will let anyone become a bio-terrorist overnight."
- The Reality: AI is smart at knowing things, but it's not great at doing things in the real world, especially for beginners.
- The Conclusion: The AI gave a modest boost (like a 1.4x improvement in getting through the steps), but it didn't turn a novice into an expert overnight. The "tacit knowledge" (the feel, the touch, the visual judgment) is still a human barrier that AI hasn't fully cracked yet.
Summary in One Sentence
Giving beginners access to super-smart AI didn't let them build a dangerous virus from scratch any more often than just using Google, but it did help them navigate the confusing middle steps a little faster, proving that knowing the answer isn't the same as having the hands-on skill to execute it.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.