科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Journal of general internal medicine2026-09-15

Optimizing Large Language Models for Hospital Discharge Prediction.

Jonathan Spagnoli, Natalie M Guzman, Karan Desai, Tony H Chang, Anthea Bell, Veronica Gilbert, Samuel Lehn, Vikas I Parekh, Michael W Sjoding, Andrew Wong

一句话结论 · In one sentence

In this retrospective cohort study, automated prompt optimization improved LLM discharge-prediction performance into the range reported for prior discharge-prediction approaches. Operational workflows, rather than medical knowledge alone, represented the majority of LLM errors.

原始摘要(英文原文)· Original abstract
OBJECTIVE: To evaluate baseline LLM performance for hospital discharge prediction, characterize LLM errors through clinically grounded qualitative analysis, and test inference-time optimization strategies to improve accuracy. MATERIALS AND METHODS: We conducted a retrospective cohort study with qualitative error analysis performed from June 2025 to December 2025 at a tertiary academic medical center. Two independent randomized cohorts of hospitalized inpatients ≥ 18 years of age admitted between January 1, 2024, and December 31, 2024, with a length of stay between 2 and 14 days were used in separate validation and test sets. LLMs predicted same-day discharge using clinical documentation from the 30 h preceding a 06:00 index time. Performance was assessed using F1 score, balanced accuracy, sensitivity, specificity, positive predictive value, and negative predictive value. Qualitative error analysis was conducted to identify LLM error domains. Three inference-time optimization strategies were tested: test-time scaling, expert-led prompt engineering, and automated prompt optimization. RESULTS: A validation set (n = 860) and test set (n = 886) were randomly generated. The baseline GPT-5 prompt achieved an F1 score of 0.48 and sensitivity of 0.37 on the validation set. Qualitative analysis identified operational workflows as the most common error source. Automated prompt optimization had higher F1 score and sensitivity than the basic prompt on the hold-out test set, with lower PPV and specificity; test-time scaling and expert-led prompt engineering showed minimal improvement. CONCLUSION: In this retrospective cohort study, automated prompt optimization improved LLM discharge-prediction performance into the range reported for prior discharge-prediction approaches. Operational workflows, rather than medical knowledge alone, represented the majority of LLM errors.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Optimizing Large Language Models for Hospital Discharge Prediction. — 科研速览 Science Skim