科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Heart & lung : the journal of critical care2026-09-16

Large language models fail to reliably predict emergent catheterization laboratory activation from prehospital electrocardiograms.

Emile Legendre, Urska Cvek, Stewart Greathouse, Colton Toups, Brandon Watkins, Dillon Jones, David Janese

一句话结论 · In one sentence

Multimodal LLM interpretation of prehospital ECGs demonstrated clinically unreliable performance for identifying ECGs warranting emergent cardiac catheterization laboratory activation when benchmarked against real-world cardiology activation decisions. Although some models achieved high sensitivity, poor specificity resulted in excessive false-positive activation recommendations. These findings suggest that general-purpose LLMs are not appropriate for ECG-based catheterization laboratory activation decisions in time-sensitive cardiopulmonary care workflows.

原始摘要(英文原文)· Original abstract
BACKGROUND: Rapid and accurate electrocardiogram (ECG) interpretation is essential for timely identification of ST-elevation myocardial infarction (STEMI) and activation of reperfusion pathways in emergency care. OBJECTIVES: To evaluate the diagnostic performance of multimodal LLMs in identifying prehospital ECGs warranting emergent catheterization laboratory activation. METHODS: We performed a retrospective analysis of 615 ECGs from 270 emergency medical service patient encounters (EMS) with concern for acute myocardial infarction. The reference standard was cardiology activation of the STEMI pathway for emergent angiography. LLM-based image interpretation (three models) and ECG machine algorithm interpretations were compared. Sensitivity, specificity, positive predictive value, negative predictive value, and overall accuracy were calculated. RESULTS: Gemini demonstrated the highest sensitivity (95.3%; 95% CI 91.7-97.3) but extremely poor specificity (9.4%), indicating a high false-positive rate. ChatGPT and Claude showed moderate sensitivity (68.1% and 67.2%) with limited specificity (42.3% and 46.5%). The ECG machine algorithm demonstrated more balanced performance, with sensitivity of 67.7% (95% CI 61.4-73.4) and higher specificity (64.2%) than all LLMs. CONCLUSIONS: Multimodal LLM interpretation of prehospital ECGs demonstrated clinically unreliable performance for identifying ECGs warranting emergent cardiac catheterization laboratory activation when benchmarked against real-world cardiology activation decisions. Although some models achieved high sensitivity, poor specificity resulted in excessive false-positive activation recommendations. These findings suggest that general-purpose LLMs are not appropriate for ECG-based catheterization laboratory activation decisions in time-sensitive cardiopulmonary care workflows.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Large language models fail to reliably predict emergent catheterization laboratory activation from prehospital electrocardiograms. — 科研速览 Science Skim