科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Cell Reports Medicine2026-02-01· Computer science

Benchmarking large language models for predictive modeling in biomedical research with a focus on reproductive health

Reuben D. Sarwal, Victor Tarca, Claire Dubin, Nikolaos Kalavros, Gaurav Bhatti, Sanchita Bhattacharya, Atul J. Butte, Roberto Romero, Gustavo Stolovitzky, Tomiko Oskotsky, Adi L. Tarca, Marina Sirota

原始摘要(英文原文)· Original abstract
Large language models (LLMs) are increasingly used for code generation and data analysis. This study assesses LLM performance across four predictive tasks from three DREAM challenges: gestational age regression from transcriptomics and DNA methylation and classification of preterm birth and early preterm birth from microbiome data. We prompt LLMs with task descriptions, data locations, and target outcomes and then run LLM-generated code to fit prediction models and determine accuracy on test sets. Among the eight LLMs tested, o3-mini-high, 4o, DeepseekR1, and Gemini 2.0 can complete at least one task. R code generation is more successful (14/16) than Python (7/16). OpenAI’s o3-mini-high outperforms others, completing 7/8 tasks. Test set performance of the top LLM-generated models matches or exceeds the median-participating team for all four tasks and surpasses the top-performing team for one task ( p = 0.02). These findings underscore the potential of LLMs to democratize predictive modeling in omics and increase research output. • Large language models (LLMs) generate R and Python code for omics prediction tasks • Top LLM matched or exceeded median DREAM challenge participant performance • LLMs allowed model development in minutes starting with tabular data • Single-shot prompting reveals strengths and limitations of LLM-based code generation Sarwal et al. benchmark large language models (LLMs) on four reproductive health prediction tasks from DREAM challenges. Several LLM-generated executable workflows match median participant accuracy and, in one case, outperform top teams. This demonstrates that LLMs can successfully assist with developing predictive machine learning models, while still requiring human oversight.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Benchmarking large language models for predictive modeling in biomedical research with a focus on reproductive health — 科研速览 Science Skim