Xiaoyan Zhou, Yue Li, Ting Ding, Jiali Liu, Dongdong Tong, Yudong Mu, Nan Xu, Sipeng Li, Hao Meng, Ning Gao, Qian He
Early diagnosis of breast cancer (BC) remains challenging. The limited sensitivity and specificity of existing serum tumor markers for reliable clinical application highlight the need to develop a more accurate and efficient screening workflow. This study analyzed serum samples from 255 breast cancer patients and 300 healthy controls using matrix-assisted laser desorption/ionization time-of-flight (MALDI-TOF) mass spectrometry, identifying 58 differentially expressed peptides (37 upregulated, 21 downregulated). Combined with machine learning, peptide identification, and external validation, a complete standardized workflow was established. Nine machine learning (ML) algorithms were employed and compared, including SVM, LightGBM, XGBoost, etc. The models were interpreted using SHAP and LIME to identify key features. Peptides of interest were sequenced via mass spectrometry. Their expression and potential prognostic value were further validated in breast cancer transcriptomic datasets. Nine machine learning algorithms showed favorable discriminatory ability in the study cohort. The LightGBM model achieved an AUC of 0.97 internally and maintained an AUC of 0.88, an accuracy of 0.8543, and a precision of 0.9799 externally. However, after correcting for the markedly elevated prevalence (80.3%) in the external cohort, the positive predictive value (PPV) decreased substantially under real-world screening scenarios, warranting prospective validation in true screening populations. Model interpretation and subsequent sequencing identified six core biomarker peptides: Apolipoprotein A-IV (APOA4), Serum Deprivation Response Protein (SDPR), Alpha-1-Antitrypsin (SERPINA1), Ezrin (EZR), Serglycin (SRGN), and Fibrinogen Alpha Chain (FGA). Transcriptomic corroboration suggested that these molecules were significantly dysregulated in breast cancer tissues and showed univariate prognostic associations with patient survival. These findings demonstrated the potential of a proteomics-driven integrated machine learning pipeline as a proof-of-concept auxiliary risk-stratification tool for enhancing early breast cancer diagnosis, warranting further prospective validation in real-world screening cohorts before clinical translation.