Shriyank Somvanshi, Anandi Dutta, Subasish Das
Abstract The growing deployment of vehicles equipped with driver assistance and automated driving systems presents new challenges for crash record classification, as existing police-reported databases vary considerably in the completeness and consistency of automation-level metadata. This study benchmarks five tabular machine learning and deep learning models: random forest, XGBoost, MambaAttention, prior-data fitted network (TabPFN), and TabTransformer for classifying reported SAE automation categories using structured crash records from the Texas CRIS database (2024), comprising 4649 records across assisted driving (SAE Level 1), partial automation (SAE Level 2), and advanced automation (SAE Levels 3–5). The SAE automation label in each record reflects the vehicle’s designed automation capability derived from make, model, and year specifications rather than confirmed system engagement at crash time; the classification task, therefore, addresses SAE-coded vehicle capability identification rather than active automation detection. Records were partitioned at the crash level using GroupShuffleSplit to prevent crash-level data leakage, yielding 3721 training and 928 test records with zero Crash_ID overlap. SMOTE-NC was applied exclusively within the training partition to address the 21:1 class imbalance, interpolating continuous features while assigning categorical values by mode selection to preserve feature integrity. Fivefold stratified cross-validation with SMOTE-NC applied independently within each fold provided robust performance estimates. XGBoost achieved the highest macro-F1 (0.573; 95% CI 0.514–0.633), while TabPFN attained the highest macro-AUC (0.786) and accuracy (82.9%) without requiring task-specific training. MambaAttention achieved a macro-F1 of 0.509 and the highest advanced automation recall among the well-performing models (27.6%; 95% CI 11.1–45.2%), reflecting a trade-off between aggregate accuracy and minority-class sensitivity; this estimate carries substantial uncertainty given only 29 advanced automation test cases. McNemar’s significance tests confirmed that TabPFN and XGBoost form a statistically equivalent performance tier ( $$p = 0.082$$ p = 0.082 ), while MambaAttention’s error profile differs significantly from all other models ( $$p < 0.001$$ p < 0.001 ). TabTransformer exhibited persistent cross-validation instability (macro-F1: $$0.212 \pm 0.100$$ 0.212 ± 0.100 ) and near-chance discriminative performance (AUC: 0.550) despite multiple architectural stabilization interventions, constituting a reproducible negative finding for vanilla transformer architectures under extreme tabular class imbalance. Feature contribution analysis using exact TreeSHAP decompositions for XGBoost identified population group, vehicle body style, roadway class, and crash timing as the dominant predictors of reported automation category. Each SAE category exhibits a distinct associative feature profile consistent with the geographic distribution and fleet composition of differently automated vehicles in the dataset. These associations are model-specific and exploratory; demographic features reflecting data collection patterns and geographic fleet distribution should not be interpreted as causal crash mechanisms. The findings demonstrate that tabular machine learning models can classify SAE-coded vehicle capabilities in police-reported crash records with meaningful accuracy, supporting applications in crash database quality assurance, automation-label inference at scale, and the design of automation-stratified safety research.