Batuhan Erdoğdu, Duygu Gül, Berrak Itır Aylı, Mine Karadeniz, Murat Yıldırım, Melda Cömert
Background: Multiple myeloma (MM) diagnosis requires specialised testing that is not universally available, contributing to diagnostic delays. The aim of this study was to develop and internally validate statistical and machine learning (ML) models that discriminate MM from clinically similar conditions using exclusively routine laboratory parameters. Methods: This retrospective diagnostic prediction model study included 422 consecutive patients referred to hematology for suspected MM at a tertiary care center (2017-2025). Fifteen routine demographic and laboratory variables were used to train seven ML algorithms and two ensemble methods. Model performance was evaluated on a held-out test set (30%) using the area under the receiver operating characteristic curve (AUC) with bootstrap 95% confidence intervals, calibration metrics, and decision curve analysis. Internal validation comprised repeated 10-fold cross-validation and bootstrap optimism correction (1000 resamples). Results: Of 422 patients, 203 (48.1%) were diagnosed with MM. Random Forest, Gradient Boosting, and Ensemble Stacking showed similarly high discrimination on the test set (AUCs 0.950-0.955). Total protein (multivariable-adjusted OR: 11.95), calcium (OR: 9.62), hemoglobin (OR: 4.34), and albumin (OR: 0.16) were the strongest independent predictors. Ensemble Stacking showed the best calibration (Brier score 0.084; calibration slope 1.14). At a ≥95% sensitivity threshold, Random Forest maintained 77.3% specificity. Risk stratification effectively separated patients, with observed MM rates ranging from 2.4% (very low risk) to 97.2% (very high risk). Conclusions: Machine learning models using routine laboratory parameters achieved excellent diagnostic performance for discriminating MM from its clinical mimickers in a real-world referral cohort. Following further development with larger datasets and external validation, these models could support triage and prioritization of patients already under evaluation for suspected MM, particularly where access to specialized testing is delayed; use in lower-prevalence, first-contact settings would require recalibration and dedicated validation.