Longmei Chen, Xiaoyan Teng, Jiale Tian, Yuzhen DU, Wanchao Liu
Routine blood biomarkers can support interpretable machine-learning models for differentiating malignant from benign breast nodules. The proposed model may serve as a complementary pre-biopsy risk-assessment tool alongside conventional imaging modalities; however, prospective validation across broader populations and healthcare settings remains necessary before routine clinical implementation.
BACKGROUND: Precise identification of malignant breast nodules is critical for clinical intervention. This study aimed to develop and externally validate an interpretable machine learning model using routine blood biomarkers.
METHODS: This retrospective multicenter study included 899 women with pathologically confirmed breast nodules in a derivation cohort from Shanghai Baoshan Hospital (March 2022-December 2024). After missing-value imputation and LASSO selection, eight models were developed. Model performance was evaluated using the area under the curve (AUC), calibration curves, and decision curve analysis (DCA). The models were subsequently validated in an independent temporal validation cohort (n = 205, recruited from the same center between January 2025 and September 2025) and an external validation cohort (n = 290, Tongji Hospital, Shanghai, between January 2025 and November 2025). Model interpretability was evaluated using SHapley Additive exPlanations (SHAP).
RESULTS: Eleven predictors (age and ten biomarkers) were selected. The Random Forest (RF) model exhibited the best performance, yielding AUCs of 0.82 (95% CI: 0.77-0.87), 0.78 (95% CI: 0.71-0.86), and 0.72 (95% CI: 0.66-0.78) in the internal, temporal, and external validations cohorts, respectively, indicating a moderate decline yet acceptable discriminative ability in independent validation cohorts. In the temporal validation cohort, its diagnostic performance was comparable to ultrasound (AUC = 0.76) and mammography (AUC = 0.74). SHAP analysis identified hs-CRP, RBC, and age as the most influential predictors. A web-based calculator was developed to facilitate future evaluation of the model.
CONCLUSION: Routine blood biomarkers can support interpretable machine-learning models for differentiating malignant from benign breast nodules. The proposed model may serve as a complementary pre-biopsy risk-assessment tool alongside conventional imaging modalities; however, prospective validation across broader populations and healthcare settings remains necessary before routine clinical implementation.