科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Cureus2025-10-13· Medical diagnosis

Piloting Temperature-Driven Variability in Emergency Diagnostic Accuracy Using a Leading Large Language Model

Philip Jarrett, Jared Hill, Marshall Howell, Kristen Moore, Joby Thoppil, Laura Vargas Ortiz, Samuel Parnell, D. Mark Courtney, Samuel McDonald, Deborah Diercks, Andrew Jamieson, Dazhe Cao

原始摘要(英文原文)· Original abstract
Background Large language models (LLMs) use a parameter known as temperature to control the stochasticity of output sampling during text outputs, which may have implications for clinical diagnostic tasks. In this study, we aimed to determine the impact of the temperature parameter on GPT-4o's diagnostic accuracy when evaluating emergency medicine cases and assess the effect on diagnostic divergence across iterations. Methodology We conducted a simulation-based diagnostic accuracy study using four challenging emergency medicine cases adapted from the Foundations of Emergency Medicine curriculum. Each case was submitted to GPT-4o 250 times at five temperature settings (0.0, 0.25, 0.50, 0.75, 1.0), both with and without physical examination findings, yielding 10,000 total outputs. Each output contained exactly three differential diagnoses with one leading diagnosis to limit the inflation of diagnostic accuracy by larger, unprioritized lists of remotely possible diagnoses. Diagnostic accuracy was assessed by comparing outputs against predetermined gold-standard diagnoses. Mixed-effects models evaluated the relationship between temperature and diagnostic accuracy, while a sensitivity analysis excluded physical examination data. Diagnostic divergence, defined as the number of unique diagnoses generated across iterations, was explored within cases as a representation of internal consistency. Results At temperature 0.0, GPT-4o achieved 100% leading diagnosis accuracy across all cases with physical examination data. As the temperature increased, the accuracy declined systematically to 89.4% at the temperature setting of 1.0. Mixed-effects models demonstrated that temperature was inversely associated with correct leading diagnosis (β = -4.16, odds ratio (OR) = 0.02, 95% confidence interval (CI) = 0.01-0.03, p < 0.001) and with inclusion of the gold-standard diagnosis anywhere in the differential (β = -3.75, OR = 0.02, 95% CI = 0.01-0.05, p < 0.001). Diagnostic divergence increased from an average of 4.5 unique diagnoses at temperature 0.0 to 26.25 at temperature 1.0 (483% increase). Case sensitivity varied significantly, with ascending cholangitis showing the greatest temperature sensitivity (accuracy dropping from 100% to 70.4%), while carbon monoxide poisoning maintained 100% accuracy across all settings. Sensitivity analysis evaluating the impact of physical examination on diagnostic accuracy revealed case-specific effects. While the diagnostic accuracies of ascending cholangitis and myxedema coma were heavily affected by the exclusion of physical examination data, the carbon monoxide and cryptococcal meningitis cases were minimally changed, if at all. Conclusions Increasing the GPT-4o temperature parameter systematically introduced diagnostic inaccuracy across four emergency medicine vignettes. Lower temperature settings led to improved diagnostic accuracy and consistency across case iterations, which may make them preferable for clinical applications requiring high reliability. Transparent reporting of temperature settings is essential for reproducible clinical artificial intelligence research.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Piloting Temperature-Driven Variability in Emergency Diagnostic Accuracy Using a Leading Large Language Model — 科研速览 Science Skim