C Y Lin
Train accident data are essential for quantitative risk analyses, and advances in data analytics have considerably improved the efficacy of data-driven risk assessments. Despite the importance of train accident data quality and resolution, limited research has focused on how accident data structures can be improved to better assist more accurate and effective risk analyses. This research presents a novel data diagnosis framework that systematically examines accident data structure and identifies key weaknesses (i.e., missing or low-resolution data fields) from a risk analysis perspective. Recommendations are made to improve train accident data structures to enhance data availability, completeness, and resolution. A case study based on the United States railway accident database was conducted to demonstrate the diagnosis framework. Results showed that information such as curvature, grade, train consist, and railcar loading status should be added to train accident data. Furthermore, the key to obtaining comprehensive and high-resolution accident data is connecting train accident data with external databases that provide important information for risk analyses such as weather conditions and hazmat through connection identifiers. This research contributes to the development of risk-oriented train accident data to enable more effective and accurate risk analysis for critical current and future rail safety topics.