Liang Chen, Jia Wei, Guoqing Wang, Xiaoxiao Yang, Lusheng Qin
Traffic accidents on highway are often characterized by high destructiveness and severe casualties. Predicting accident severity and understanding its causes are crucial for enhancing highway safety. To address the issues of limited prediction accuracy and poor interpretability of traditional machine learning and deep learning methods at the current stage, this study proposes an accident severity prediction model based on a hybrid architecture of MobileNetV3 and a Transformer. The model first encodes numerical accident-related variables into two-dimensional images using the Gramian Angular Field (GAF) method. Local spatial features are then extracted via the depthwise separable convolution modules of MobileNetV3, and long-range temporal dependencies are captured through the Transformer encoder, which outputs the final prediction. The proposed model is compared with Convolutional Neural Networks (CNNs), Long Short-Term Memory networks (LSTMs), MobileNetV3, a Transformer, and LSTM–Transformer architectures in terms of prediction performance. Results show that the MobileNetV3–Transformer model achieves the highest accuracy of 0.9549. Finally, the DeepSHAP interpretability algorithm is introduced to reveal the systemic influence and contribution of significant factors to accident severity. The results indicate that vehicle age, special road conditions, speed limits, and lighting conditions are closely related to the severity of highway accidents. This study provides a reliable theoretical basis for early warning of highway accidents and refines control measures to further enhance highway safety.