AssayBench: An Assay-Level Virtual Cell Benchmark for LLMs and Agents
This paper introduces AssayBench, a new benchmark comprising 1,920 CRISPR screens designed to evaluate the ability of large language models and agents to predict cellular phenotypic outcomes, revealing that zero-shot generalist LLMs currently outperform specialized biology models while highlighting significant room for improvement toward achieving robust virtual cell models.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer
Imagine you are a detective trying to solve a massive biological mystery. In a real lab, scientists run thousands of experiments where they "turn off" specific genes in cells to see what happens. Does the cell stop growing? Does it get sick? Does it become resistant to a drug? These experiments are called CRISPR screens, and they generate huge lists of genes ranked by how important they are to the outcome.
The problem is that running these experiments in real life is slow, expensive, and messy. Scientists wish they had a "Virtual Cell"—a super-smart computer program that could predict the results of these experiments before they ever touch a petri dish.
This paper introduces AssayBench, a new "final exam" designed to test if our current Artificial Intelligence (AI) models are smart enough to be that Virtual Cell.
Here is how the paper breaks it down, using simple analogies:
1. The New Exam: AssayBench
Before this paper, AI models were mostly tested on narrow tasks, like predicting how a single molecule changes shape. But real drug discovery is more like a movie plot: you need to understand the characters (genes), the setting (the cell type), and the plot twist (the drug treatment) to guess the ending (the phenotype).
The authors built AssayBench using 1,920 real-world experiments that have already been published.
- The Task: The AI is given a text description of an experiment (e.g., "We turned off genes in leukemia cells and added a drug called Etoposide to see which cells survived").
- The Goal: The AI must output a ranked list of the top 100 genes that it thinks caused the effect, from "most likely" to "least likely."
- The Twist: The AI has to do this for new experiments it hasn't seen before, just like a student taking a test on a topic they studied but haven't practiced on specifically.
2. The Grading System: The "Adjusted Score"
How do you grade an AI on a list of 1,000 genes? If you just ask, "Did it get the top one right?" it's too harsh. If you ask, "Did it get any right?" it's too easy.
The authors created a special grading metric called Adjusted nDCG. Think of it like a golf handicap:
- Some experiments are easy (like a flat golf course); some are hard (like a course with sand traps).
- The metric adjusts the score so that an AI gets credit for beating the "random guess" baseline, but it doesn't get a perfect score just because the experiment was easy.
- It also penalizes the AI if it puts "bad" genes (genes that actually make the problem worse) at the top of the list.
3. The Contestants: Who Played?
The authors tested three types of "students":
- The General Geniuses (Frontier LLMs): These are massive, general-purpose AI models (like Gemini 3 Pro or GPT-5.4) that read almost everything on the internet. They weren't trained specifically for biology; they just know a lot about everything.
- The Biology Specialists: These are AI models trained specifically on biological data or designed to act as medical agents.
- The "Cheat Sheet" Baselines: Simple methods that just look for patterns, like "Which genes are usually important in this type of experiment?"
4. The Results: The Generalists Won (For Now)
Surprisingly, the General Geniuses (the big, general AI models) performed the best.
- The Specialist models actually did worse than the general ones.
- The "Cheat Sheet" baselines were surprisingly competitive, especially for common cell types, suggesting that sometimes simple patterns are enough.
- The Verdict: Even the best AI models are still far from perfect. The paper shows that the current best AI scores are roughly halfway to what is theoretically possible (the "performance ceiling"). If you took two real scientists and asked them to predict the same experiment, they would likely agree much more than the AI does.
5. The "Memorization" Trap
The authors noticed something interesting: The AI did much better on older experiments (published before 2021) than on very recent ones (published in late 2025).
- The Analogy: It's like a student who memorized the answers to last year's practice tests but struggles with the new questions on today's exam.
- The Conclusion: The AI is likely "memorizing" facts it read in its training data rather than truly understanding the biological logic. However, it still performed better than specialists on the new tests, suggesting it has some genuine reasoning ability, not just memory.
6. How to Get Better?
The paper tested a few ways to boost the AI's performance:
- Fine-tuning: Teaching the AI specifically on these types of problems helped a little.
- Ensembling: Asking three different AIs to vote on the answer and combining their results worked the best. This is like asking a panel of experts rather than just one.
The Bottom Line
AssayBench is a new, realistic test to see if AI can predict how cells will behave when we mess with their genes.
- Current Status: AI is getting good at this, but it's not there yet. The best models are still making mistakes that real scientists wouldn't.
- The Future: The paper suggests that to build a true "Virtual Cell," we need models that can reason through biology rather than just memorizing facts, and we need more data to teach them.
The paper does not claim that these models are ready to be used in hospitals or to cure diseases today. It simply says: "Here is a ruler to measure how far we have to go before AI can truly replace the lab bench."
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.