科研速览 · Science Skim继续刷下去 · Keep skimming →
◇ bioRxiv2026-08-22· bioinformatics

LifeSciBench: Evaluating Language Models on Realistic, Expert-Level Tasks in the Life Sciences

A. Liu, A. Ho, A. M. Droste, D. Martin, E. Wong, E. Zhou, I. Zhou, J. Park, J. Jiao, K.-R. Skelly, K. Kim, J. Li, K. Rao, M. Uehara, M. Marion, N. Fitzgerald, R. Dias, S. Shringarpure, Y. Yuan, Y. Wang

原始摘要(英文原文)· Original abstract
We introduce LifeSciBench, a benchmark of 750 expert-authored tasks designed to evaluate whether language models can handle realistic life science research work. The majority of existing life sciences benchmarks have a narrow scope or are purely knowledge-based, and therefore fail to capture the complexity of real-world research, which often involves ambiguities and requires the accurate execution of multiple dependent judgment calls. Additionally, almost all existing benchmarks span at best a small collection of subdomains within the life sciences; there is at present no existing life sciences benchmark with both the requisite breadth and depth required to convincingly measure proficiency in real-world professional research settings. LifeSciBench addresses this gap by spanning seven representative scientific workflows and seven life science domains, with each constituent task paired with a human expert-written rubric. Across five frontier and domain-specialized models, GPT-Rosalind performs best, with a task-weighted mean normalized rubric score of 0.576 and a task-weighted response pass rate of 36.1% (response-level values are first averaged within each task, and the resulting task-level values are then averaged with equal weight). LifeSciBench remains unsaturated, with 171 tasks (22.8%) having no observed passing response from any evaluated model and 261 tasks (34.8%) having a best-model pass rate below 20%. LifeSciBench therefore serves as a high-resolution evaluation of practical scientific reasoning and operational decision-making in the life sciences.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

LifeSciBench: Evaluating Language Models on Realistic, Expert-Level Tasks in the Life Sciences — 科研速览 Science Skim