D. Weissenbacher, J. Vo, A. Uy-Evanado, K. Reinier, H. Chugh, V. Kadiyala, K. OCOnnor, S. S. Chugh, G. Gonzalez-Hernandez
Background: Identification of in-hospital cardiac arrest (IHCA) through manual chart abstraction is the gold standard, but its time-consuming and resource-intensive, which naturally limits its practicality. Diagnostic codes provide an accessible automated alternative, but prior work has shown this method of extraction has poor sensitivity and poor positive predictive value. More accurate and scalable automated methods are needed to support IHCA surveillance and research. Methods: We developed large language model (LLM)-based systems that incorporated physician-defined clinical knowledge through guideline-enriched prompting, supervised fine-tuning, and a self-correction procedure. The systems identified IHCA events and their locations from electronic health record notes. Performance was evaluated at the encounter level against physician-adjudicated chart review and compared with ICD-code-based identification. Results: Compared with an ICD-code-based algorithm (precision 0.68, recall 0.91, F1 score 0.78), the supervised fine-tuned LLM offered substantially better performance for IHCA identification on a curated evaluation corpus, reaching a precision of 0.88, a recall of 1.00, and an F1 score of 0.94. When tested on a larger, unselected validation cohort of 45,525 encounters, a separate LLM configuration - relying on guideline-enriched prompting and a self-correction procedure rather than fine-tuning - achieved a precision of 0.79 for encounter-level IHCA identification, meaning most encounters flagged by the model were genuine IHCA cases. The fine-tuned LLM underperformed on this cohort due to its inability to self-correct its decisions, an issue that can be mitigated by having a separate model perform the correction step. Conclusion: Explicit integration of clinical knowledge through prompting and supervised fine-tuning enabled accurate identification of IHCA events and locations from routine clinical documentation. In a two-pass workflow, these systems could prioritize encounters for clinician adjudication, reduce manual screening requirements, and support more scalable IHCA surveillance and registry development.