科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Journal of the American Medical Informatics Association : JAMIA2026-09-08

Evaluating retrieval-augmented generation versus long-context input for clinical reasoning over electronic health records.

Skatje Myers, Dmitriy Dligach, Timothy A Miller, Samantha Barr, James Landefeld, Yanjun Gao, Matthew M Churpek, Anoop Mayampurath, Majid Afshar

一句话结论 · In one sentence

Our results suggest that RAG remains a competitive and efficient approach for clinical tasks over large amounts of EHR, even as newer models become capable of handling increasingly longer amounts of text.

原始摘要(英文原文)· Original abstract
OBJECTIVE: To evaluate whether retrieval-augmented generation (RAG) can serve as an efficient alternative to long-context prompting for clinical reasoning over electronic health records (EHRs). MATERIALS AND METHODS: We defined 3 EHR-based tasks that are replicable across health systems and vary in reasoning complexity: (1) extracting imaging procedures (modality, date, and anatomic site), (2) generating timelines of therapeutic antibiotic use, and (3) identifying the key diagnoses for a hospitalization. Using real inpatient clinical notes from a US academic health system, we evaluated 3 large language models (GPT-5.4-mini, Mistral Medium 3, DeepSeek V3.1) with varying amounts of provided context, comparing targeted retrieval to using the most recent clinical notes. RESULTS: For Imaging Procedures, RAG strongly outperformed recent-note inputs and exceeded long-context performance (by 0.17-9.83 F1 across all models) using fewer than 8K tokens. Similar benefits were observed for Antibiotic Timelines, where <8K of retrieved tokens matched long-context recent-notes performance (between -3.26 and +3.24 Jaccard). Error analysis revealed that missing information in the clinical notes-often due to inter-hospital transfers-limited performance to some extent. However, performance on the Diagnosis Generation task remains largely static across methods and models. DISCUSSION: RAG demonstrated strong token efficiency across tasks, with the clearest and most consistent gains observed for imaging extraction and antibiotic timeline reconstruction. Diagnosis generation proved the most challenging task, suggesting ceiling effects imposed by documentation variability and evaluation constraints. CONCLUSION: Our results suggest that RAG remains a competitive and efficient approach for clinical tasks over large amounts of EHR, even as newer models become capable of handling increasingly longer amounts of text.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Evaluating retrieval-augmented generation versus long-context input for clinical reasoning over electronic health records. — 科研速览 Science Skim