← Latest papers
🤖 AI

TxBench-PP: Analyzing AI Agent Performance on Small-Molecule Preclinical Pharmacology

This paper introduces TxBench-PP, a verifiable benchmark for evaluating AI agents on small-molecule preclinical pharmacology tasks using real-world assay data, revealing that even the most advanced current models fail to reliably recover accurate preclinical decisions.

Original authors: Hannah Le, Ramesh Ramasamy, Alex Urrutia, Mahsa Yazdani, Tim Proctor, Kenny Workman

Published 2026-06-19
📖 4 min read☕ Coffee break read

Original authors: Hannah Le, Ramesh Ramasamy, Alex Urrutia, Mahsa Yazdani, Tim Proctor, Kenny Workman

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are hiring a team of super-smart, AI-powered research assistants to help you decide which new medicines to build. You have a massive pile of messy, real-world lab data—charts, microscope photos, and chemical test results—and you need these assistants to look at the evidence and tell you: "This drug looks promising, let's keep it," or "This one is dangerous or useless, let's throw it out."

This paper, TxBench-PP, is a report card on how well these AI assistants actually performed that specific job.

The Big Idea: A "Driver's Test" for AI

The authors created a special test called TxBench-PP. Think of it like a driving test, but instead of a car, the AI is driving a drug discovery program.

  • The Trap: They didn't just ask the AI to recite facts from a textbook (like "What does aspirin do?"). That's easy for AI because they memorized the internet.
  • The Real Challenge: They gave the AI a fresh, messy folder of new data from a real lab experiment and asked, "Based only on these specific numbers and images, what should we do?"
  • The Goal: To see if the AI can act like a scientist who looks at the evidence, or if it just guesses based on what it remembers from its training.

The Results: The AI is Still a Rookie

The researchers tested 16 different combinations of AI models (the "brains") and software tools (the "hands"). They ran the test 4,800 times in total.

Here is the verdict: The AI is not ready to drive alone yet.

  • The Best Score: The top-performing AI system got about 59% of the answers right. In a classroom, that's a failing grade. In a real drug company, getting it wrong 4 out of 10 times is dangerous because it could mean wasting millions of dollars or, worse, moving a dangerous drug forward.
  • The Reality Check: Even the "smartest" AI failed to get the right answer on the same task three times in a row. It was inconsistent.

Where Did the AI Go Wrong?

The authors looked closely at the mistakes, and they found the AI wasn't just "forgetting" things. It was making specific types of scientific errors:

  1. The "Textbook" Trap: Sometimes the AI saw a picture of a cell dying and ignored the actual data in front of it. Instead, it remembered a famous story from a textbook about a different drug and applied that story here. It ignored the evidence to fit a pre-written narrative.
  2. Bad Math and Bad Filters: The AI often messed up the "quality control." Imagine a scientist looking at a graph and saying, "Oh, that weird spike is just a glitch, I'll ignore it," when actually, that spike was the most important part of the data. The AI frequently threw away important data or kept garbage data.
  3. The "All-or-Nothing" Problem: When asked to pick the best drug from a list of 10, the AI often picked the one that looked "okay" but missed the one that was truly great, or it picked a dangerous one because it looked active. It struggled to make the hard choice of "kill this project" vs. "advance this project."

The "Tool" Matters

One interesting finding was that the software the AI used to do the work mattered just as much as the AI itself.

  • Think of the AI as a brilliant chef.
  • Harness A gave the chef a sharp knife and a clean cutting board.
  • Harness B gave the same chef a dull butter knife and a dirty board.
  • The chef with the good tools (called "Pi" in the paper) cooked much better meals than the same chef with the bad tools, even though the "chef" (the AI model) was identical.

The Bottom Line

This paper is a reality check. While AI promises to speed up drug discovery, it cannot yet be trusted to make the critical "go/no-go" decisions on its own.

The current AI agents are like interns who are very good at reading the manual but still need a senior scientist to look over their shoulder, check their math, and make sure they aren't ignoring the messy reality of the lab data. They are helpful, but they are not yet reliable enough to run the show.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →