科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ BMJ digital health & AI2026-01-01

COPE: Chain-of-Thought Prediction Engine for open-source large language model-based stroke outcome prediction from clinical notes.

Yongkai Liu, Helena Feng, Bin Jiang, Yixin Wang, Max Wintermark, David Liebeskind, Michael Moseley, Maarten Lansberg, Gregory Albers, Jeremy Heit, Greg Zaharchuk

一句话结论 · In one sentence

COPE, a reasoning-enhanced framework using lightweight, open-source LLMs, achieved performance comparable to a proprietary model and to strong traditional baselines while operating without model retraining or manual feature engineering. It offers an accurate and privacy-preserving solution for outcome prediction from unstructured clinical text.

原始摘要(英文原文)· Original abstract
OBJECTIVE: To develop and evaluate Chain-of-Thought Outcome Prediction Engine (COPE), a reasoning-enhanced large language model, for predicting 90-day functional outcomes after acute ischaemic stroke (AIS) from unstructured clinical notes. METHODS: We included 464 patients with AIS who had discharge summaries and 90-day modified Rankin Scale (mRS) outcomes. COPE uses a two-step chain-of-thought (CoT) framework based on sequential open-source models (LLaMA-3-8B): the first generates intermediate clinical reasoning and the second outputs an mRS prediction. We compared COPE's performance with GPT-4.1, ClinicalBERT, a structured variable-based machine learning model (XGBoost) and a single-step large language model (LLM) without CoT. Performance was evaluated using mean absolute error (MAE), accuracy within ±1 mRS point (±1 ACC) and exact accuracy (ACC). RESULTS: COPE achieved an MAE of 1.01 (95% CI 0.92 to 1.11), ±1 ACC of 74.4% (95% CI 69.9% to 78.8%) and ACC of 32.8% (95% CI 28.0% to 37.6%), comparable to GPT-4.1 (MAE, 1.00 (95% CI 0.90 to 1.09); ±1 ACC, 77.9% (95% CI 73.7% to 82.0%); ACC, 32.5% (95% CI 28.0% to 37.4%); p=0.72, 0.11 and 0.96). COPE demonstrated performance comparable to a strong structured-data baseline using XGBoost (MAE, 1.03 (95% CI 0.93 to 1.13); ±1 ACC, 73.9% (95% CI 69.4% to 78.5%); ACC, 33.3% (95% CI 28.5% to 38.2%); p=0.77, 0.87 and 0.89), and outperformed ClinicalBERT and the single-step LLM. Subgroup analyses showed consistent performance across sex and age, with higher error among older patients, those undergoing thrombectomy and those with longer summaries. CONCLUSIONS: COPE, a reasoning-enhanced framework using lightweight, open-source LLMs, achieved performance comparable to a proprietary model and to strong traditional baselines while operating without model retraining or manual feature engineering. It offers an accurate and privacy-preserving solution for outcome prediction from unstructured clinical text.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

COPE: Chain-of-Thought Prediction Engine for open-source large language model-based stroke outcome prediction from clinical notes. — 科研速览 Science Skim