Anna Erickson, L. R. Jackson, Shreya Chappidi, Ethan Kahana, Kevin Camphausen, Andra V. Krauze
Abstract Introduction Primary CNS tumors impact 25,000 individuals annually in the U.S., posing significant health challenges. Research is hindered by fragmented data across electronic health record (EHR) systems and the inefficiency of manual extraction for diagnostic, molecular, and treatment data. Large Language Models (LLMs) present a solution by automating extraction, enhancing efficiency and consistency, and enabling AI-driven insights. Methods We analyzed 4,974 EHR documents from 256 patients with confirmed glioblastoma (GBM). All patients were treated on NCI NIH IRB (IRB00011862)-approved protocols 00-C-0074, 02C0064, 04C0200, 06C0112, 16-C-0081, and 20-C-0027. Cohort 1 (GBM, n = 109) was used to develop to the pipeline, while Cohorts 2 (various CNS histologies, n = 147) and 3 (GBM lesions, n = 15) served as testing sets. Clinical features extracted with GPT-4o via the NIH NIDAP Text Extraction Program (NTEP) included date of diagnosis, KPS, extent of resection, MGMT and IDH statuses, and radiation therapy start/end dates. Prompts underwent iterations, utilized JavaScript Object Notation (JSON) formatting, and outputs were then compared to the manual ground truth, established through detailed chart review of all 256 patients across three cohorts to extract key clinical variables of interest, allowing for the analysis of clinical feature and document type accuracy. Results Prompt refinement led to a 30-fold increase in prompt character count, achieving ≥ 95% accuracy for five of seven features in Cohort 1. MGMT accuracy increased from 26% to 99%, and radiation dates increased from 93% to 98%. KPS (82%) and extent of resection (84%) were less accurate. For Cohort 2, MGMT and IDH were extracted with the highest accuracy (87% and 90%, respectively), while Cohort 3 achieved slightly lower but still strong performance (70% and 78%). Across these cohorts, MGMT and IDH were the most accurately extracted biomarkers. Radiation therapy summaries were the most effective documents across all three cohorts based on the extraction rates of clinical features. Conclusion LLMs can accurately extract clinical features from EHRs with careful document selection and prompt design. This method may support CNS tumor research and have broader clinical applications. Future work will expand feature sets and validate on external datasets.