Duo Dai, Min Ma
BackgroundConventional breast cancer screening is limited by cost, access, and patient discomfort. Whether a lightweight machine-learning model trained on routine laboratory indicators can support opportunistic screening remains unsettled.MethodsA retrospective cohort of 13,285 individuals (1,216 ICD-10 C50 cases, 12,069 controls) from a Chinese tertiary hospital during 2022 to 2024 was analysed with a leakage-controlled pipeline. Six classical learners and LightGBM were trained on twenty indicators, then assessed by discrimination, calibration, decision-curve analysis, and zero-shot UK Biobank validation.ResultsThe lightweight LightGBM achieved an internal AUC of 0.987 with sensitivity 0.835, specificity 0.983, Brier 0.026, and an external AUC of 0.83, training in 5.6 s within 1.65 MB.ConclusionThe model is a candidate low-cost triage tool, pending multicentre prospective validation and recalibration.