科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Multiple sclerosis and related disorders2026-09-23

Generative AI responses to exercise-related questions asked by patients with multiple sclerosis: An evaluation of quality, accuracy, and readability.

Aylin Aydogdu Delibay, Cimen Olcay Demir, Nisa Turutgen, Humeyra Kiloatar, Simge Donmez

一句话结论 · In one sentence

Within the scope of this study, ChatGPT-generated responses demonstrated higher quality, accuracy, and reliability than those generated by Gemini. Nonetheless, patient accessibility could be limited by the poor readability metrics observed in both tools. While AI-based chatbots may show promise in reinforcing patient education for MS populations, they must not substitute specialized medical professionals during clinical decision-making.

原始摘要(英文原文)· Original abstract
OBJECTIVE: This study aims to compare the quality, accuracy, reliability, and readability of responses produced by ChatGPT-5.4 Thinking and Google Gemini 3 Flash regarding frequently asked questions by multiple sclerosis (MS) patients about exercise. METHOD: A total of 75 questions were evaluated. Expert physiotherapists analysed the responses utilizing the Global Quality Score (GQS), Modified DISCERN (mDISCERN), Likert Accuracy Scale, and the Flesch Reading Ease (FRE). RESULTS: While 61.3% of ChatGPT responses were classified as high quality, this rate was 24% for Gemini (p < 0.001). ChatGPT-generated responses demonstrated significantly higher overall accuracy and mDISCERN scores than those generated by Gemini (p < 0.001). Categorical analyses revealed that the responses generated by ChatGPT received higher accuracy in the domains of exercise planning, safety and risk management, as well as participation and self-management. Additionally, ChatGPT-generated responses received significantly higher mDISCERN scores in the exercise planning and safety categories. No significant difference was observed between the two models regarding overall FRE scores (p > 0.05). The mean FRE scores for ChatGPT and Gemini were 41.70 and 38.21, respectively, with responses from both models classified at a "difficult" readability level. CONCLUSION: Within the scope of this study, ChatGPT-generated responses demonstrated higher quality, accuracy, and reliability than those generated by Gemini. Nonetheless, patient accessibility could be limited by the poor readability metrics observed in both tools. While AI-based chatbots may show promise in reinforcing patient education for MS populations, they must not substitute specialized medical professionals during clinical decision-making.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Generative AI responses to exercise-related questions asked by patients with multiple sclerosis: An evaluation of quality, accuracy, and readability. — 科研速览 Science Skim