科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Studies in health technology and informatics2026-09-17

Benchmarking AI Vibe Coding for Clinical Statistical Analysis: A Structured Evaluation in Pulmonary Hypertension Research.

Jeeva Sam, T Spreuer, M Berger, A Yogeswaran, T Khodr, P Janetzko, W Seeger, J Bienzeisler, R W Majeed

原始摘要(英文原文)· Original abstract
Statistical analysis of clinical data requires expertise in medical statistics. Large language models (LLMs) are increasingly used for code generation and may support both descriptive and advanced analyses, but their reliability remains uncertain. This study evaluated whether five current LLMs (GPT 5.3, Claude Sonnet 4.6, Gemini 2.5 Flash, Perplexity, and Grok) could reproduce expert-validated statistical analyses from a published clinical workflow across three statistical tasks: descriptive table generation, Kaplan-Meier survival analysis, and Cox proportional hazards modelling. All models received identical datasets and standardised prompts. Their outputs were compared with analyses performed by two experts trained in mathematical statistics and evaluated for grouping correctness, numerical accuracy, missing-value handling, and quality of generated R code. All five models produced correct descriptive statistics once dataset variables were specified explicitly. Two models failed the initial descriptive benchmark because of variable-name ambiguity, but both recovered after prompt clarification. All models reproduced the correct Kaplan-Meier p-value, although figure completeness differed. In the Cox benchmark, four models reproduced all required hazard-ratio terms, while one omitted the interaction results. These findings suggest that LLMs may support clinical research workflows, but their outputs still require careful validation before use.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Benchmarking AI Vibe Coding for Clinical Statistical Analysis: A Structured Evaluation in Pulmonary Hypertension Research. — 科研速览 Science Skim