← Latest papers
🤖 AI

BioSecBench-Surveillance: A Verifiable Benchmark for AI Agents in Pathogen Genomic Surveillance

The paper introduces BioSecBench-Surveillance, a verifiable benchmark demonstrating that even the most advanced AI agents currently struggle to reliably infer correct genomic analysis pipelines for pathogen surveillance, achieving only around 50% accuracy across diverse tasks.

Original authors: Harmon Bhasin, Kevin Flyangolts, Dianzhuo Wang, Evan Seeyave, Arjun Banerjee, Amanda Darling, Joshua Stallings, David Stern, Shawn Higdon, Claire Duvallet, Bryan Tegomoh, Kenny Workman

Published 2026-07-22
📖 4 min read☕ Coffee break read

Original authors: Harmon Bhasin, Kevin Flyangolts, Dianzhuo Wang, Evan Seeyave, Arjun Banerjee, Amanda Darling, Joshua Stallings, David Stern, Shawn Higdon, Claire Duvallet, Bryan Tegomoh, Kenny Workman

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine the world of public health as a massive, high-stakes detective agency. Their job is to spot invisible villains—viruses, bacteria, and other pathogens—before they can start a global party crash, like a pandemic. To do this, scientists use a super-powerful microscope called "genomic surveillance." Instead of just looking at a germ under a lens, they read its entire instruction manual (its DNA or RNA) to figure out exactly what it is, where it came from, and how dangerous it might be.

For a long time, the hard part was just getting enough of these instruction manuals. But now, machines are so good at reading them that we are drowning in data. The real bottleneck has shifted: it's no longer about finding the clues; it's about solving the mystery. This is where Artificial Intelligence (AI) agents come in. Think of these agents as super-smart, tireless junior detectives. They can read millions of pages in seconds. But here's the big question: Can these AI detectives actually solve the case, or do they just get lost in the library? Can they look at a messy pile of genetic clues and decide, "Okay, I need to use this specific tool with these specific settings to find the answer," just like a human expert would?

This paper, BioSecBench-Surveillance, is like a giant, rigorous final exam for these AI detectives. The researchers created a "verifiable benchmark"—a standardized test with 100 different scenarios—where an AI agent is given raw genetic data and a specific surveillance context (like "this is wastewater from a city" or "this is a sample from a hospital"). The AI has to figure out the right analysis pipeline from scratch: which databases to check, which filters to apply, and which thresholds to set, all without being told the answer.

The results of this exam were a bit of a reality check. Out of 16 different combinations of AI models and software tools tested, the very best ones only got about 50.2% of the questions right. That means even the smartest AI agents failed half the time. The top performers were a model called Opus 4.8 paired with a tool called PI, and another called GPT-5.5 with Codex, both scoring 50.2%. They were followed closely by Opus 4.7 at 49.6% and Sonnet 4.6 at 48.6%. The weakest configurations, like some versions of Grok, only managed to get about 14% to 16% of the tasks correct.

The paper found that the AI agents weren't usually failing because they didn't know which tools to use. In fact, they almost always picked the right "detective tools" (like the correct software for reading DNA). Their mistakes happened in the fine print: they chose the wrong reference books, set the wrong sensitivity thresholds, or messed up the math when normalizing the data. It's like a detective who knows they need a fingerprint scanner but sets the sensitivity so high that they miss the print, or so low that they think a smudge is a match.

The difficulty of the test varied wildly depending on the type of case. The easiest tasks for the AI were source tracking (figuring out where a sample came from) and taxonomic classification (naming the organism), where they got about 50% and 46% right, respectively. However, the AI struggled mightily with anomaly detection (spotting weird, unexpected signals), where they only got 20% right. They also had a hard time with genetic-engineering characterization (detecting if a virus was artificially modified), scoring only 35%. Interestingly, the type of sample (like wastewater vs. clinical samples) didn't matter as much as the technology used to read the DNA. Long-read sequencing data was much harder for the AI to handle (26% pass rate) compared to short-read data (41%).

The researchers also noticed that some AI models were too shy. About 27% to 31% of the time, certain models simply refused to answer the question, saying they couldn't do it, whereas others refused almost never. This "refusal" rate varied wildly depending on which software "harness" the model was running on.

Ultimately, the paper concludes that while AI agents are getting better, they are not yet reliable enough to be trusted as the sole decision-makers in pathogen surveillance. They can run the analysis, but they often make the wrong judgment calls about how to run it. The authors suggest that until AI can reliably make these nuanced choices—like knowing exactly which reference to trust or which threshold to set—we can't fully rely on them to stop the next outbreak. For now, the human expert is still the one holding the final say on the case.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →