Qiongge Li, Jian Liu, Xing Li, Wei Nie, Jiajin Fan
LLM-based classification demonstrates strong potential for automating retrospective safety event tagging and streamlining future reporting workflows. In a proposed future system, staff would only need to provide a detailed narrative description while the AI model performs classification and prompts for human validation when uncertainty arises. This hybrid approach can improve efficiency, consistency, and scalability in radiation oncology safety reporting. To our knowledge, this is the first study to apply an LLM for automated classification of radiation oncology safety events, introducing a proof-of-concept framework with potential for broader applicability for AI-assisted incident reporting and quality improvement.
PURPOSE: This study aims to evaluate the feasibility of using a large language model (LLM) to automate the classification of radiation oncology safety events and to outline a framework for future reporting systems that leverage artificial intelligence (AI)-assisted workflows.
METHODS AND MATERIALS: Retrospective safety event data were extracted from our institutional reporting system and processed using Python scripts for deidentification and formatting. GPT-5 (OpenAI) was accessed via an application programming interface and refined through iterative prompt engineering (a natural language process that does not modify model parameters) to classify incidents across multiple dimensions, including failure mode, severity, treatment type, discoverer's role, exclusion criteria, and occurring/discovering workflow stages. The model's performance was validated by three independent, blinded expert reviewers, with inter-rater agreement quantified using Cohen's κ.
RESULTS: On a blinded 80-incident validation set, the model's classification fell within the experts' group-accepted answer in 67.9% of dimension-level comparisons on average (96.2% for the inclusion decision; 85.3% for treatment type and 76.6% for occurred workflow). Agreement between the model and individual reviewers (mean Cohen's κ = 0.31) was comparable to agreement among the reviewers themselves (mean κ = 0.42), indicating performance approaching that of an independent expert. Severity scoring showed the greatest variability for both the model and the human reviewers (model-reviewer κ = 0.14; inter-reviewer κ = 0.22), highlighting it as the primary area for improvement. The comparable model-expert and expert-expert agreement supported the application of the model to full data set analysis without further optimization.
CONCLUSIONS: LLM-based classification demonstrates strong potential for automating retrospective safety event tagging and streamlining future reporting workflows. In a proposed future system, staff would only need to provide a detailed narrative description while the AI model performs classification and prompts for human validation when uncertainty arises. This hybrid approach can improve efficiency, consistency, and scalability in radiation oncology safety reporting. To our knowledge, this is the first study to apply an LLM for automated classification of radiation oncology safety events, introducing a proof-of-concept framework with potential for broader applicability for AI-assisted incident reporting and quality improvement.