D. Wang, Q. Xu, A. Banerjee, R. Jain, A. Vermani, J. Felix, G. Moreno, F. AlZaben, N. Keul, E. Seeyave, H. Bhasin
Inferring biological function from experimental data is central to understanding emerging pathogens and developing effective countermeasures, yet interpreting these data remains slow and expert-intensive. AI agents could help accelerate this process by reasoning across sequence, structural, and biophysical evidence. We present BioSecBench-Function, a verifiable benchmark for recovering biosecurity-relevant function from real biological data. The benchmark comprises 111 evaluations built from published datasets and graded deterministically against ground truth. We organize evaluations along two dimensions: threat axis (spanning seven biosecurity-relevant question types) and biological question (indicating whether the solution depends primarily on sequence, structure, or biophysical assay data). Across 7,326 runs from twenty-two model-harness configurations, Opus 5 under Claude Code led on endpoint pass rate at 50.3%, and Grok 4.6 under Grok Build led on overall pass rate at 44.1% when refusals counted as failures. Performance varied substantially across both model-harness configurations and task categories. Refusal rates differed sharply by provider, and cost was a poor predictor of accuracy: several configurations exceeded 40% pass rate at low cost. BioSecBench-Function provides a standard for measuring whether agents can be trusted to interpret what a new pathogen or variant does when the next outbreak arrives.