Zhen Hao, Jidong Yang
• Introduces CrashSage, an LLM framework for interpretable crash analysis. • Fine-tunes LLaMA3-8B, outperforming state-of-the-art tabular models and baseline LLMs. • Gradient-based attribution reveals word-level crash severity risk factors. • Co-occurrence analysis uncovers interdependent patterns among safety dimensions. Road crashes claim over 1.3 million lives annually worldwide and incur global economic losses exceeding $1.8 trillion. Such profound societal and financial impacts underscore the urgent need for road safety research that uncovers crash mechanisms and delivers actionable insights. Conventional statistical models, machine learning along with modern deep learning models, typically rely on structured crash data, overlooking contextual nuances, such as narrative elements related to multi-vehicle interactions, crash progression, and rare event characteristics, and struggling to capture complex interactions and underlying semantics. This study presents CrashSage, a novel Large Language Model (LLM)-centered framework designed to transcend these limitations and advance traffic safety analysis and modeling. Particularly, CrashSage introduces four key innovations to transform raw data into actionable intelligence. First, it pioneers a novel tabular-to-text transformation strategy with a relational data integration schema, enabling the conversion of raw, heterogeneous crash data into enriched, structured textual narratives that retain essential structural and relational contexts. Second, it employs context-aware data augmentation using a base LLM model to enhance narrative coherence while preserving factual integrity. Third, we demonstrate the power of domain specialization by fine-tuning a LLaMA3-8B model for crash severity inference. The resulting model delivers superior performance against a comprehensive suite of baselines, including state-of-the-art tabular models (CatBoost, TabTransformer, FT-Transformer) and various prompting strategies with popular LLMs (GPT-4o, GPT-4o-mini, LLaMA3-70B). Finally, and most importantly, CrashSage integrates a novel gradient-based explainability technique that illuminates model decisions at the word level for individual crashes in an intuitive manner. This technique also moves beyond identifying isolated risk factors to uncovering their complex interplay. It reveals how driver behaviors act as central nodes that connect with environmental and infrastructure factors and highlights dangerous synergistic effects. By emulating the analytical workflow of a human expert, CrashSage establishes a new paradigm for traffic safety analysis to offer interpretable insights into the complex dynamics of road crashes.