L. K. Chinthala, C. Lemon, A. Shaban-Nejad, G. Farage, R. L. Davis, H. Xu, C. Madlock-Brown
Background: Social isolation and social support are clinically important but inconsistently represented in structured electronic health record data. Their identification from clinical narratives is complicated by nuanced, context-dependent language and by variation in how social connectedness is documented across clinical settings, patient populations, and geographic locations. Models intended for health care use must therefore distinguish social isolation from documented support, negation, ambiguous social circumstances, and non-social clinical uses of similar terminology while remaining robust to differences in documentation practices. Objective: This study aimed to develop and evaluate locally deployable, fine-tuned language models for context-aware classification of social isolation, social support, and contextually similar nonrelevant clinical text across multiple health systems. Materials and Methods: We conducted a retrospective study using electronic health record data from 326,847 adults aged 50 years or older who received care at 3 Tennessee health systems from 2020 to 2023. Candidate clinical text was identified using social-environment ICD-10-CM Z codes and an established social isolation lexicon. Human annotators classified 9,748 spans from 9,578 clinical notes contributed by 6,881 patients as social isolation, social support, or no social isolation. The annotated corpus included social isolation and support concepts including negations, infection-control and microbiology terminology, institutional living arrangements, and other contextually ambiguous examples. Uncertainty-based active learning was used to prioritize informative spans for annotation. FLAN-T5-Large, BERT, RoBERTa, and Gemma-2-2B were fine-tuned for 3-class span classification. Full fine-tuning and parameter-efficient fine-tuning using low-rank adaptation approaches were evaluated. Performance was assessed using site-based internal-external cross-validation, with macro-F1 as the primary metric. Results: The annotated corpus included 2,318 social isolation spans, 638 social support spans, and 6,792 spans classified as no social isolation. Fully fine-tuned FLAN-T5-Large achieved the highest and most consistent performance, with a mean macro-F1 of 0.92 (SD 0.04). Class-specific F1 scores were 0.91 (SD 0.03) for social isolation, 0.90 (SD 0.04) for social support, and 0.94 (SD 0.05) for no social isolation. Gemma-2-2B achieved a macro-F1 of 0.89 (SD 0.10), compared with 0.80 (SD 0.21) for RoBERTa and 0.77 (SD 0.17) for BERT. Full fine-tuning also outperformed parameter-efficient fine-tuning in the corresponding held-out evaluation. Remaining errors primarily involved patients who lived alone but received assistance, the presence of family members without meaningful support, grief-related language, and non-social clinical uses of isolation terminology. Conclusions: Fine-tuned language models can distinguish social isolation, social support, and contextually similar but unrelated language across heterogeneous clinical settings. FLAN-T5-Large provided the strongest and most stable performance, including for the less frequent social support class. The 3-class annotation framework, inclusion of difficult negative examples, and multisite evaluation offer a practical approach for converting nuanced social-connectedness documentation into structured data. Further work should evaluate end-to-end identification in unselected clinical notes and prospective integration into clinical and population health workflows.