科研速览 · Science Skim继续刷下去 · Keep skimming →
◇ bioRxiv2026-09-08· biophysics

BioSecBench-Function: A Verifiable Benchmark for Reasoning about Biological Function from Experimental Data

D. Wang, Q. Xu, A. Banerjee, R. Jain, A. Vermani, J. Felix, G. Moreno, F. AlZaben, N. Keul, E. Seeyave, H. Bhasin

原始摘要(英文原文)· Original abstract
Inferring biological function from experimental data is central to understanding emerging pathogens and developing effective countermeasures, yet interpreting these data remains slow and expert-intensive. AI agents could help accelerate this process by reasoning across sequence, structural, and biophysical evidence. We present BioSecBench-Function, a verifiable benchmark for recovering biosecurity-relevant function from real biological data. The benchmark comprises 111 evaluations built from published datasets and graded deterministically against ground truth. We organize evaluations along two dimensions: threat axis (spanning seven biosecurity-relevant question types) and biological question (indicating whether the solution depends primarily on sequence, structure, or biophysical assay data). Across 7,326 runs from twenty-two model-harness configurations, Opus 5 under Claude Code led on endpoint pass rate at 50.3%, and Grok 4.6 under Grok Build led on overall pass rate at 44.1% when refusals counted as failures. Performance varied substantially across both model-harness configurations and task categories. Refusal rates differed sharply by provider, and cost was a poor predictor of accuracy: several configurations exceeded 40% pass rate at low cost. BioSecBench-Function provides a standard for measuring whether agents can be trusted to interpret what a new pathogen or variant does when the next outbreak arrives.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

BioSecBench-Function: A Verifiable Benchmark for Reasoning about Biological Function from Experimental Data — 科研速览 Science Skim