Jin Baek Kwon
Student dropout undermines both student success and institutional performance. In recent years, various machine-learning models have been developed to predict dropout risk in higher education; however, existing studies remain constrained by four major gaps: (i) limited prediction horizons, (ii) first-year focus, (iii) inadequate handling of class imbalance, and (iv) lack of calibrated interpretability. To overcome these limitations, this study proposes a portable, generalizable machine-learning framework for long-term dropout prediction that systematically integrates modular components addressing each gap. Specifically, (i) rolling multi-semester windows generate horizon-specific targets several semesters in advance, (ii) an all-cohort student–semester dataset eliminates the first-year restriction by representing every enrolled cohort, (iii) a resampling–threshold optimization coupling mitigates class imbalance to improve sensitivity to rare events, and (iv) probability calibration with SHAP (SHapley Additive exPlanations)-based interpretability ensures reliable and transparent model outputs. Empirical evaluation on a university registry dataset demonstrates that the framework produces well-calibrated, temporally stable, and interpretable dropout-risk estimates. Because it relies solely on ubiquitous registry variables and off-the-shelf algorithms, the framework is inherently portable and generalizable across higher-education institutions and suitable for integration into early-warning systems that support timely, data-driven student advising.