科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Journal of imaging informatics in medicine2026-09-18

Large Language Models for Ankle Fracture Classification and Management Prediction from Routine Clinical Documentation: A Single-Center Exploratory Study.

Benjamin Schwarberg, Conrad Ketzer, Bosse Thiel, Andreas Sauter, Marcus R Makowski, Daniel Spitzl, Markus Mergen, Florian T Gassert

原始摘要(英文原文)· Original abstract
Ankle fractures are among the most common injuries in trauma surgery and require accurate classification, consistent documentation, and individualized management. Large language models (LLMs) offer the potential to translate unstructured clinical text reports into structured, management-related information, yet their role in orthopedic trauma workflows remains insufficiently defined. In this retrospective study, the performance of Llama 3.1 (70B Instruct) was evaluated using routine radiology reports and clinical documentation from 54 patients with acute ankle fractures. Outputs were compared with reference standards derived from routine clinical documentation for Weber fracture classification, operative versus nonoperative management, and surgical procedure category. Four prompting strategies were systematically assessed. Accuracy for Weber classification ranged from 0.759 to 0.815, well above majority-class baseline (0.556), with strong performance for Weber B fractures and most errors occurring between adjacent categories. Macro-averaged F1 across the three Weber classes ranged from 0.744 to 0.822. Operative management prediction reached accuracies of 0.833 to 0.963 (sensitivity 0.872 to 1.000, specificity 0.571 to 0.714). Procedure category prediction demonstrated lower accuracy (0.426 to 0.610 depending on the endpoint), largely at or below a majority-class baseline (0.553), reflecting the limitations of text-only input for operative planning. These findings suggest that LLMs can extract structured information from routine clinical documentation, performing well for standardized classification but less reliably for procedure-level prediction. Prospective and multimodal validation is needed before clinical implementation.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Large Language Models for Ankle Fracture Classification and Management Prediction from Routine Clinical Documentation: A Single-Center Exploratory Study. — 科研速览 Science Skim