BioSecBench-Function: A Verifiable Benchmark for Reasoning about Biological Function from Experimental Data
The paper introduces BioSecBench-Function, a verifiable benchmark comprising 111 evaluations across seven biosecurity-relevant question types to assess the ability of AI agents to accurately infer biological functions from experimental data, revealing significant performance variations among models and highlighting that cost is not a reliable predictor of accuracy.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer
Technical Summary: BioSecBench-Function
Problem Statement
Inferring biological function from experimental data is a critical bottleneck in understanding emerging pathogens and developing countermeasures. Currently, this interpretive process is characterized as slow and heavily dependent on human expertise. While AI agents hold the potential to accelerate this workflow by reasoning across diverse evidence types—including sequence, structural, and biophysical data—there is a lack of standardized, verifiable frameworks to assess whether these agents can reliably recover biosecurity-relevant functions from real-world biological datasets.
Methodology
To address this gap, the authors introduce BioSecBench-Function, a verifiable benchmark designed to evaluate AI agents on the task of recovering biological function from experimental data. The benchmark is constructed from published datasets and utilizes deterministic grading against ground truth to ensure objective evaluation.
The evaluation framework is organized along two primary dimensions:
- Threat Axis: Comprises seven distinct question types relevant to biosecurity.
- Biological Question: Categorizes tasks based on the primary data modality required for the solution, specifically distinguishing between dependencies on sequence data, structural data, or biophysical assay data.
The benchmark consists of 111 evaluations. The study conducted 7,326 runs across twenty-two different model-harness configurations to test performance variability.
Key Results
The evaluation yielded the following performance metrics and observations:
- Top Performers: When refusals were counted as failures, Opus 5 (under the Claude Code harness) achieved the highest endpoint pass rate at 50.3%. Grok 4.6 (under the Grok Build harness) achieved the highest overall pass rate at 44.1%.
- Performance Variability: There was substantial variation in performance across different model-harness configurations and specific task categories.
- Refusal Rates: Refusal rates differed sharply depending on the AI provider.
- Cost vs. Accuracy: The study found that cost was a poor predictor of accuracy. Notably, several configurations achieved pass rates exceeding 40% at low costs.
Significance and Claims
The paper positions BioSecBench-Function as a necessary standard for measuring the reliability of AI agents in high-stakes biological contexts. Its primary significance lies in providing a mechanism to determine whether agents can be trusted to interpret the functional implications of new pathogens or variants during future outbreaks. By establishing a verifiable benchmark grounded in real biological data and ground truth, the work aims to move beyond theoretical capabilities to practical, measurable trustworthiness in biosecurity applications.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.