Haoyu Yang, Chi‐Yang Li, Qingsheng Wang
Process industry incidents can inflict fatalities, heavy economic losses, and reputational harm, making robust Process Safety Management (PSM) essential. Yet the volume and free-text nature of modern investigation records outstrip human analytic capacity. We present a domain-augmented large language model (LLM) framework that combines structured Chain-of-Thought prompting with retrieval-augmented generation (RAG) to classify offshore-platform incident reports into the 11 top-level root cause categories defined in the ABS Group Root Cause Map™. A total of 1182 Bureau of Safety and Environmental Enforcement (BSEE) district investigation reports were parsed into structured JSON format, while ABS definitions and illustrative examples were embedded into a FAISS index to support on-the-fly retrieval. Four prompting setups were benchmarked with GPT-4o-mini on 100 randomly selected reports. The Domain CoT + RAG model raised document-level set-based F1-score from 0.552 (zero-shot baseline) to 0.663, driven by a 17 % higher precision. Predicted categories per report fell from 3.92 to 2.78, close to the human average of 2.55, showing that domain context curbs over-classification without sacrificing coverage. Category-level analysis revealed high performance in major categories such as Equipment Reliability, Procedure , and Human Factors (F1 > 0.75), while challenges persisted in Design and Documentation , which require more implicit causal reasoning. Conditional-probability mapping across the full dataset reproduced expected clusters of human-related failures consistent with risk-based PSM theory. These findings demonstrate that combining domain-specific prompts with information retrieval significantly enhances the reasoning capacity of generative LLMs for multi-label safety analytics, offering a scalable, low-cost pathway toward digitizing incident investigations. • An LLM-based framework is proposed to perform automated classification of offshore incidents. • Domain-specific CoT and RAG proves effective in enhancing the overall performance. • Overclassification cut significantly by embedding RCA guidelines and examples. • LLM reproduces human-related root cause clusters predicted by RBP. • The framework could offer efficient and scalable route to digitized RCA in PSM.