Qianying He, Xuan Liu, Jingquan Liu, Bao Liu, Wenjian Liu
Procedural abstraction accounted for a larger share of the observed gain than graph propagation, while external miscalibration and dataset shift limited transportability of absolute risk. ProMem-Agent should therefore be interpreted as a retrospective research framework for studying reusable clinical-trajectory representations, not as a clinically deployable decision-support system. Prospective, site-specific, clinician-in-the-loop evaluation remains necessary.
BACKGROUND AND OBJECTIVE: Clinical large language model (LLM) agents can interpret current clinical context but have limited mechanisms for converting longitudinal experience into compact, reusable units. We developed ProMem-Agent, a procedural-memory framework that represents recurrent early intensive-care trajectories as provenance-linked observational patterns for retrospective mortality-risk estimation.
METHODS: Adult ICU stays with at least 24 hours of observable data were represented as six consecutive four-hour state-action-response intervals. Memory candidates were extracted exclusively from the MIMIC-IV ICU training cohort, linked to source events, consolidated by semantic clustering, and organized in a similarity graph. For each new patient, hybrid semantic and graph retrieval selected three memory cards. Comparators included conventional and longitudinal EHR models, direct and Chain-of-Thought LLM prompting, patient-level Case-RAG, token-matched Case-RAG, semantic-only Procedure-RAG, and first-24-hour SOFA as a clinically established severity reference. Evaluation included patient-level bootstrap testing, calibration analysis, external validation, retrieval and clustering sensitivity analyses, perturbation experiments, and blinded expert review.
RESULTS: The internal test cohort contained 6368 ICU stays with 11.9% mortality. ProMem-Agent achieved an F1-score of 0.608, AUC of 0.836, AUPRC of 0.481, and Brier score of 0.086. Relative to Procedure-RAG, the incremental differences were modest (AUC +0.013; AUPRC +0.029). Without memory reconstruction or external recalibration, AUC/AUPRC values were 0.823/0.402 on eICU, 0.831/0.429 on MIMIC-III, and 0.803/0.361 on HiRID. External calibration slopes were 0.88, 0.91, and 0.85, respectively, compared with 0.97 internally.
CONCLUSIONS: Procedural abstraction accounted for a larger share of the observed gain than graph propagation, while external miscalibration and dataset shift limited transportability of absolute risk. ProMem-Agent should therefore be interpreted as a retrospective research framework for studying reusable clinical-trajectory representations, not as a clinically deployable decision-support system. Prospective, site-specific, clinician-in-the-loop evaluation remains necessary.