Sunyang Fu, Min Ji Kwak, Jaerong Ahn, Qiuhao Lu, Zhiyi Yue, Nan Wang, Huipeng Liu, Grace Giles, Alexa Cumming, Keziah M Thomas, Shreyans Sanghvi, Jiang Jun, Ming Huang, Xiaoyang Ruan, Andrew Wen, Liwei Wang, Erin Hommel, Yanshan Wang, Lichao Sun, Huiwen Xu, Chan Mi Park, Jennifer St Sauver, Jeffrey S Wefel, Nahid Rianon, Dae Hyun Kim, Hongfang Liu
Current Natural Language Processing (NLP) algorithms for detecting geriatric conditions are largely limited to domain-specific models that fail to capture the interdependent, multidimensional nature of comprehensive geriatric assessment. This study aimed to develop and evaluate a comprehensive, scalable, and robust information extraction framework to identify Comprehensive Geriatric Assessment (CGA) and Age-Friendly Health Systems (AFHS) 4Ms-related data elements from unstructured electronic health record (EHR) text across multiple health systems. Using a team science approach grounded in the TRUST framework, we annotated pooled clinical notes from four health systems to produce a gold-standard dataset of 41 CGA- and 4Ms-related geriatric care data elements. Three information extraction approaches were implemented and evaluated: an in-context learning generative large language model (GPT-4o), a hybrid heuristic-LLM model (MedAgingIE), and an instruction-tuned open-source lightweight model (Qwen2-7B-Instruct). Performance was assessed on a blinded test set using macro- and micro-averaged metrics. GPT-4o achieved a macro F1-score of 0.56 and micro F1-score of 0.87; MedAgingIE achieved 0.55 and 0.92; and Qwen2-7B-Instruct achieved 0.30 and 0.81, respectively. MedAgingIE demonstrated the strongest consistency between precision and recall, while GPT-4o showed superior sensitivity for diverse, context-rich geriatric concepts. These findings highlight key trade-offs among symbolic, generative, and instruction-tuned approaches for CGA and 4Ms phenotyping, suggesting that hybrid heuristic-LLM methods offer interpretability and stability, whereas large language models provide greater adaptability for complex clinical narratives.