Tahera Alnaseri, Yasmine Ibrahim, Ansgar Grunseid, Mehrnaz Siavoshi, Alaina Matthews, Clifford Y Ko, Michael R DeLong
In this proof-of-concept validation study, a customized LLM achieved significantly higher abstraction accuracy than conventional human review for general and breast reconstruction NSQIP variables. These findings support LLM-assisted abstraction as a promising approach to improve efficiency, reduce resource requirements, and facilitate broader implementation and expansion of NSQIP, although multicenter validation remains necessary.
BACKGROUND: National Surgical Quality Improvement Program (NSQIP) data collection depends on labor-intensive manual chart abstraction, limiting efficiency, increasing cost, and necessitating patient sampling. This study evaluated whether a large language model (LLM) could accurately abstract unstructured NSQIP breast reconstruction variables compared with conventional human abstraction.
STUDY DESIGN: Clinical notes from patients enrolled in the NSQIP Breast Reconstruction pilot program (July 1, 2024-February 28, 2025) were manually de-identified and processed using a customized ChatGPT 4.1 workflow targeting individual variables. A faculty plastic surgeon established the reference standard. Overall accuracy of LLM and human abstraction was compared using McNemar's and Chi-square tests.
RESULTS: Among 105 patients (73 bilateral, 32 unilateral), 9,048 data points were evaluated. Overall abstraction accuracy was 99.33% (61 errors) for the LLM versus 98.19% (164 errors) for human abstraction (McNemar p<0.001; Chi-square p<0.001). LLM performance exceeded human abstraction for operative and postoperative variables but was slightly lower for preoperative variables. The most frequent LLM error involved prior breast surgical history (29/61 errors), followed by prepectoral versus subpectoral implant or expander placement, a variable frequently requiring inference from documentation.
CONCLUSIONS: In this proof-of-concept validation study, a customized LLM achieved significantly higher abstraction accuracy than conventional human review for general and breast reconstruction NSQIP variables. These findings support LLM-assisted abstraction as a promising approach to improve efficiency, reduce resource requirements, and facilitate broader implementation and expansion of NSQIP, although multicenter validation remains necessary.