科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Scientific Reports2025-11-21· Data extraction

Evaluating the reliability of large language models for clinical data extraction in bladder cancer prognosis

D. H. Sun, Lubomir M. Hadjiiski, Grace Bruno, John Gormley, Heang‐Ping Chan, Elaine M. Caoili, Richard H. Cohan, Ajjai Alva, Rada Mihalcea, Chuan Zhou, Vikas Gulani

原始摘要(英文原文)· Original abstract
Advances in natural language processing (NLP) and machine learning could assist human users in clinical data extraction from unstructured electronic medical records (EMRs). This study investigates the accuracy and consistency of several Large Language Models (LLMs) - including Dolly, Vicuna, Llama, and GPT-4 - in extracting critical clinical information pertinent to bladder cancer survival prediction. Using EMRs from 163 bladder cancer patients, we assessed the impact on LLM performance by factors such as differences in the trained models, model evolution, input text length, and sequencing of case inputs. GPT-4 demonstrated superior performance with Fleiss' Kappa values exceeding 0.97, accuracy consistently above 93%, and survival prediction metrics closely aligned with ground truth (AUC ± 0.02). Among offline models, Llama-2.0-13b and Llama-3.3-70b exhibited the highest reliability in both information extraction and survival prediction. This study underscores the potential of LLMs to automate clinical data extraction for predictive modeling while highlighting the challenges related to LLM variability and reliability.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Evaluating the reliability of large language models for clinical data extraction in bladder cancer prognosis — 科研速览 Science Skim