科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ JMIR Medical Informatics2026-07-31· Interpretability

Prediction of Postoperative Vomiting Within 24 Hours Using Machine Learning With Large Language Model–Enhanced Interpretability: Development and Validation Study

Huan-Jun Wang, Wei‐Po Lee, Tz‐Ping Gau, Kuang‐I Cheng, Cheng-Ru Wei

原始摘要(英文原文)· Original abstract
Background: Postoperative nausea and vomiting are common complications after anesthesia. However, vomiting represents a clinically distinct and objectively measurable endpoint. Objective: This study aimed to develop and internally validate predictive models for postoperative vomiting within 24 hours using structured perioperative data and unstructured clinical text, while introducing a structured framework that separates feature construction from interpretability using large language models (LLMs). Methods: We analyzed 33,460 anesthesia records from a single center (2019-2022). Two temporally defined prediction tasks were constructed to reflect real-world clinical decision-making and prevent information leakage: a preoperative model using variables available before anesthesia induction, and a perioperative model using variables available up to the end of surgery. Structured data were modeled using machine learning algorithms (logistic regression, Extreme Gradient Boosting, Light Gradient Boosting Machine [LightGBM]). Unstructured clinical text was incorporated through a deterministic, concept-driven preprocessing pipeline, where LLMs were used solely for normalization (temperature=0) without feature generation, followed by rule-based concept mapping and feature encoding. Post hoc interpretability was further supported using an LLM-based Question Answering Chain module. Model performance was evaluated using receiver operating characteristic-area under the curve (AUC), precision-recall AUC, calibration metrics, and threshold-based operating characteristics. Classification thresholds were selected using the Youden J statistic, and all metrics were reported with 95% CIs derived from bootstrap resampling. Decision curve analysis was performed to assess clinical utility. Results: A total of 33,460 surgical procedures were included, of which 3607 (10.8%) experienced postoperative vomiting within 24 hours. In the preoperative task, LightGBM achieved an AUC of 0.729 (95% CI 0.706-0.749), compared with 0.610 (95% CI 0.588-0.632) for the Apfel score. In the end-of-surgery task, LightGBM achieved an AUC of 0.735 (95% CI 0.714-0.757). At the Youden-optimal threshold, the negative predictive value exceeded 0.95 across all models. Decision curve analysis demonstrated positive net benefit across clinically relevant threshold probabilities. Incorporating text-derived features provided modest improvements, while LLM-based explanation modules generated structured, natural-language explanations intended to enhance interpretability without substantially improving predictive performance. Conclusions: Machine learning models can effectively predict postoperative vomiting within 24 hours using perioperative data. The proposed framework demonstrates that LLMs can be integrated in a controlled and reproducible manner-restricted to deterministic normalization and post hoc reasoning-to generate natural-language explanations intended to enhance the interpretability of model predictions, without introducing information leakage or altering predictive modeling. As no formal clinician-based evaluation was conducted, this interpretability benefit cannot yet be objectively confirmed, and the generated explanations should be regarded as a useful interpretability aid to be validated in future clinician-centered studies. External, multicenter validation is required before broader clinical applicability can be assumed.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Prediction of Postoperative Vomiting Within 24 Hours Using Machine Learning With Large Language Model–Enhanced Interpretability: Development and Validation Study — 科研速览 Science Skim