科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Digital health2026-01-01

Evaluation and comparison of large language model responses to common discharge questions from patients with acute coronary syndrome after percutaneous coronary intervention: An expert-rated comparative study.

Dachang Qiu, Qingjiang Wang, Wenlong Yan, Sumin Yang

一句话结论 · In one sentence

Both models produced few expert-adjudicated safety flags in this limited question set. ChatGPT received higher ratings in selected quality domains, but sparse safety events and restricted question coverage preclude conclusions of equivalence or clinical safety. LLM-generated discharge information should supplement, not replace, individualized clinical advice.

原始摘要(英文原文)· Original abstract
BACKGROUND: After percutaneous coronary intervention (PCI), patients with acute coronary syndrome (ACS) require long-term medication, risk-factor control, rehabilitation, and symptom recognition. Large language models (LLMs) are increasingly used for health consultation, but their quality, safety, and stability in post-PCI discharge education remain unclear. METHODS: Twenty-five standardized Chinese discharge questions were submitted to ChatGPT 5.5 Instant and DeepSeek V3. Each model generated three responses per question (150 total). Three cardiovascular experts rated the original Chinese responses across five Likert domains and two binary safety judgments. Comparisons used exact paired permutation tests, question-clustered bootstrap confidence intervals, and Holm adjustment for six quality-score comparisons. RESULTS: Overall, 149/150 responses (99.3%; 95% CI, 96.3%-100.0%) met the primary safety criterion: 75/75 for ChatGPT (100.0%; 95% CI, 95.2%-100.0%) and 74/75 for DeepSeek (98.7%; 95% CI, 92.8%-100.0%). The paired risk difference was 1.3 percentage points (95% CI, -1.3 to 2.7; P=1.000). One DeepSeek response contained both a major error and a potentially harmful recommendation. After pairing and Holm adjustment, ChatGPT scored higher in accuracy, guideline concordance, completeness, and total quality (adjusted P=0.015, 0.015, 0.019, and 0.005). Five-domain satisfactory response rates did not differ significantly (100.0% vs 90.7%; P=0.125). CONCLUSIONS: Both models produced few expert-adjudicated safety flags in this limited question set. ChatGPT received higher ratings in selected quality domains, but sparse safety events and restricted question coverage preclude conclusions of equivalence or clinical safety. LLM-generated discharge information should supplement, not replace, individualized clinical advice.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Evaluation and comparison of large language model responses to common discharge questions from patients with acute coronary syndrome after percutaneous coronary intervention: An expert-rated comparative study. — 科研速览 Science Skim