Charlotta Lindvall, Andrea Liebowitz, Kaitlin Emmert-Tangredi, Narda J Martin, Shelby Brage, Michael K Paasche-Orlow, Yuchiao Chang, Joshua R Lakin, Jennifer Itty, Shreya Sanghani, Megan Carroll, Trey Feldeisen, Seth Randa, Aretha Delight Davis, James A Tulsky, Angelo E Volandes, Edith A Burns
BackgroundEvidence that advance care planning (ACP) improves goal-concordant care (GCC) remains limited. Challenges are lack of reproducible operational definition of GCC and labor-intensive nature of manual chart review.ObjectivesDevelop a structured approach to adjudicating GCC and compare large language model (LLM) and human assessments of concordance between goals of care (GOC) and end-of-life treatment.MethodsNatural language processing-assisted chart review created longitudinal summaries of GOC and end-of-life treatment for decedents enrolled in two pragmatic ACP trials. First, human and LLM developed a shared adjudication framework. Second, two human adjudicators and two LLMs independently assigned GCC scores and rationales for 120 purposively selected cases. Pairwise agreement quantified with Gwet's AC2 after categorizing scores as low, intermediate, or high concordance.ResultsIn phase 1, two LLMs produced near-identical scores across 11 complex cases, with human ratings within 1-2 points of the LLMs in 8 cases. In Phase 2, pairwise Gwet's AC2 coefficients ranged 0.43-0.75; agreement was highest for GPT-4o vs Human 1 (0.75) and GPT-4o vs Claude 3.5 Sonnet (0.71). Human 1 tended to assign higher and Human 2 lower scores relative to overall mean. Disagreement was greatest with changing goals, surrogate conflict, delayed code-status updates, or aggressive treatment despite documented treatment limitations.ConclusionThis proof-of-concept study demonstrates that LLMs generated GCC ratings and rationales broadly aligned with human assessments while revealing inherent subjectivity of adjudicating longitudinal goal-concordance. LLM-assisted adjudication may support scalable and standardized evaluation of documented GCC in pragmatic trials, with human review remaining essential for complex/ambiguous cases.