科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Journal of safety research2026-09-01

Improving crash data quality with large language models: Evidence from secondary crash narratives in Kentucky.

Xu Zhang, Mei Chen

一句话结论 · In one sentence

For agencies with labeled training data, fine-tuned RoBERTa is the recommended deployment choice, which offers the strongest accuracy at negligible computational cost. For agencies lacking labeled data, zero-shot LLMs such as Llama3:70B provide a viable alternative that can be deployed readily and simultaneously accumulate a labeled dataset for eventual transition to fine-tuned models.

原始摘要(英文原文)· Original abstract
INTRODUCTION: High-quality crash data are essential for traffic safety analysis, yet police-reported crash databases often suffer from underreporting and miscoding, particularly for secondary crashes. This study evaluates advanced natural language processing (NLP) techniques to enhance crash data quality by mining crash narratives, using secondary crash identification in Kentucky as a case study. METHOD: Drawing from 16,656 manually reviewed narratives from 2015 to 2022, with 3803 confirmed secondary crashes, we systematically compared 11 models across four paradigms: zero-shot open-source large language models (LLMs), fine-tuned transformers, deep learning model with word embeddings, and logistic regression. Statistical significance was assessed using pairwise McNemar's tests, and 95% bootstrap confidence intervals were computed for all metrics. RESULTS: Fine-tuned transformers achieved statistically superior performance, forming a top-performing cluster that was indistinguishable internally. RoBERTa yielded the highest F1 (0.90) and accuracy (95.4%) while requiring only seconds of inference on the test set. Among zero-shot LLMs, Llama3:70B reached the best F1 (0.86) but required 139 min of inference. The BiLSTM baseline (F1: 0.79) was statistically indistinguishable from Qwen3:32B and Gemma3:27B, while logistic baseline lagged well behind (F1: 0.66). Qualitative error analysis revealed that RoBERTa and Llama3:70B exhibit complementary failure patterns, supporting ensemble deployment strategy. CONCLUSIONS: For agencies with labeled training data, fine-tuned RoBERTa is the recommended deployment choice, which offers the strongest accuracy at negligible computational cost. For agencies lacking labeled data, zero-shot LLMs such as Llama3:70B provide a viable alternative that can be deployed readily and simultaneously accumulate a labeled dataset for eventual transition to fine-tuned models. PRACTICAL APPLICATIONS: These findings allow transportation agencies to automate labor-intensive narrative reviews, addressing chronic data quality issues like secondary crash miscoding. Practical deployment considerations are discussed, which emphasize privacy-preserving local deployment, ensemble approaches for improved accuracy, and incremental processing for scalability, providing a replicable scheme for enhancing crash-data quality with advanced NLP.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Improving crash data quality with large language models: Evidence from secondary crash narratives in Kentucky. — 科研速览 Science Skim