Nabeel Abdulrazaq Yaseen, Rana Sami Hameed, Mohammed Mahdi Hashim, Mustafa Sabah Taha
Traffic accidents remain a major challenge for urban planning and transportation engineering, posing continuous risks to human life and infrastructure.Accurately predicting accident severity-particularly the rare serious and fatal outcomes that matter most for safety policy-is difficult because such data are severely imbalanced.This study presents a leakage-controlled comparative benchmark and a set of integrated models for predicting accident severity as slight, serious, or fatal.Using a large-scale UK Department for Transport dataset (2,047,256 accident-level records joined to 2,058,408 vehicle-level rows, from which a stratified sample of 200,000 rows is modelled), we construct a leakage-safe pipeline with a temporal train/validation/test split.We evaluate logistic regression, random forest, balanced random forest, XGBoost, LightGBM, CatBoost, HistGradientBoosting, and a feed-forward neural network, each combined with imbalance strategies (class weighting, SMOTE, ADASYN, SMOTE-Tomek, focal loss, and threshold moving).We further implement three integrations: out-of-fold stacking, a two-stage cascade, and a costsensitive soft-voting ensemble.Because missing a fatal or serious case greatly exceeds the cost of a slight misclassification, the primary objectives are safety-related cost and minority-class recall.Overall accuracy is shown to be misleading: a standard random forest attains 84% accuracy yet only 1.5% fatal-class recall.Balanced random forest achieves the highest fatal-class recall (0.64), while stacking raises fatal-class recall from 0.015 to 0.527 and attains the highest balanced accuracy among integrations (0.512).A paired bootstrap (1000 resamples) confirms the macro-F1 gains of the cost-sensitive ensemble and LightGBM.Minority-class detection, while improved, remains the central open problem.