科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Frontiers in public health2026-01-01

A comparative evaluation of generative AI chatbots for patient-oriented advice on painful diabetic peripheral neuropathy: safety, accuracy, guideline concordance, actionability, and readability.

Lina Gu, Zhaole Gong, Xiaoli Qian, Ling Miao, Zhengfeng Gu

一句话结论 · In one sentence

The five chatbots showed distinct performance patterns across content quality and readability. Although no overall safety difference was detected, every system generated at least one response with a plausible pathway to inappropriate self-management, delayed assessment, medication or product misuse, or preventable injury. Chatbots may support general patient education, but medication decisions, foot-risk assessment, and urgent-care triage require professional verification.

原始摘要(英文原文)· Original abstract
BACKGROUND: Patients with painful diabetic peripheral neuropathy (PDPN) increasingly use generative artificial intelligence chatbots for information on symptoms, treatment, foot care, and when to seek professional help. Their usefulness depends on safety, accuracy, guideline concordance, actionability, and readability. OBJECTIVE: To compare five publicly accessible generative AI chatbots in answering standardized patient-oriented questions about PDPN. Methods: Sixty standardized English-language questions covering eight clinical domains were submitted once to ChatGPT, Gemini, Microsoft Copilot, DeepSeek, and Doubao in separate single-turn conversations, yielding 300 responses. Five reviewers independently assessed safety, accuracy, guideline concordance, and actionability using predefined criteria. Guideline concordance was scored against six mapped elements per question and converted to a percentage. Actionability was assessed using seven binary criteria with prespecified question-level applicability. Readability was evaluated using six established indices. Paired comparisons used Cochran's Q test for safety and Friedman tests for non-binary outcomes, followed by multiplicity-adjusted pairwise analyses. RESULTS: All 300 responses were analyzed. Inter-rater agreement was high for safety (Fleiss' κ = 0.874), accuracy [ICC (2,1) = 0.881], guideline concordance [ICC (2,1) = 0.874], and actionability [ICC (2,1) = 0.877]. Twenty-eight responses (9.3%) were classified as unsafe or potentially unsafe. Unsafe-response rates ranged from 5.0 to 15.0%, with no detected overall between-model difference (Cochran's Q = 4.462, p = 0.347). Accuracy, guideline concordance, and actionability differed across models (all p < 0.001; Kendall's W = 0.683, 0.730, and 0.556, respectively). ChatGPT generally achieved higher content-related scores, whereas Doubao scored lower. All readability indices also differed across models (all p < 0.001), with ChatGPT and Doubao producing less complex text and DeepSeek showing greater reading difficulty. CONCLUSION: The five chatbots showed distinct performance patterns across content quality and readability. Although no overall safety difference was detected, every system generated at least one response with a plausible pathway to inappropriate self-management, delayed assessment, medication or product misuse, or preventable injury. Chatbots may support general patient education, but medication decisions, foot-risk assessment, and urgent-care triage require professional verification.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

A comparative evaluation of generative AI chatbots for patient-oriented advice on painful diabetic peripheral neuropathy: safety, accuracy, guideline concordance, actionability, and readability. — 科研速览 Science Skim