Ning Du, Mengmeng Zhang, Xue Wang, Lijie Zhang, Yao Zhou, Chunxia Song
Existing predictive models exhibit considerable variability in reported discriminatory performance, with AUCs ranging from 0.644 to 0.914. However, the majority of these models are hampered by issues such as single-center development, insufficient external validation, incomplete calibration reporting, and methodological biases. The substantial heterogeneity observed across the reviewed studies precludes the generation of a single, reliable summary estimate of model performance. To enhance the clinical applicability and practical value of these predictive tools, future research should prioritize multicenter prospective studies, optimize variable handling and model validation strategies, and ensure comprehensive reporting of both discrimination and calibration.
OBJECTIVE: This systematic review aims to comprehensively map, critically appraise, and synthesize the quality and performance of existing prediction models for prolonged air leak (PAL) following pulmonary resection, with a focus on their clinical utility and potential for translation.
METHODS: A systematic search was conducted in CNKI, Wanfang Data, VIP, SinoMed, PubMed, Web of Science, Embase, and the Cochrane Library, up to March 2026. Two researchers independently screened the literature, extracted data, and assessed the risk of bias in the predictive models. A qualitative, narrative synthesis of model characteristics and performance was performed. Results: A total of 26 models were included. Most studies had good applicability, but the overall risk of bias was high. The models showed a wide range of discriminatory power, with Area Under the Curve(AUC)values ranging from 0.644 to 0.914. Frequently identified predictors included pleural adhesions, forced expiratory volume in 1 second(FEV1), body mass index(BMI), age, smoking history, and gender.
CONCLUSION: Existing predictive models exhibit considerable variability in reported discriminatory performance, with AUCs ranging from 0.644 to 0.914. However, the majority of these models are hampered by issues such as single-center development, insufficient external validation, incomplete calibration reporting, and methodological biases. The substantial heterogeneity observed across the reviewed studies precludes the generation of a single, reliable summary estimate of model performance. To enhance the clinical applicability and practical value of these predictive tools, future research should prioritize multicenter prospective studies, optimize variable handling and model validation strategies, and ensure comprehensive reporting of both discrimination and calibration.
SYSTEMATIC REVIEW REGISTRATON: https://www.crd.york.ac.uk/PROSPERO/home, identifier CRD420261435579.