Aylin Aydogdu Delibay, Cimen Olcay Demir, Nisa Turutgen, Humeyra Kiloatar, Simge Donmez
Within the scope of this study, ChatGPT-generated responses demonstrated higher quality, accuracy, and reliability than those generated by Gemini. Nonetheless, patient accessibility could be limited by the poor readability metrics observed in both tools. While AI-based chatbots may show promise in reinforcing patient education for MS populations, they must not substitute specialized medical professionals during clinical decision-making.
OBJECTIVE: This study aims to compare the quality, accuracy, reliability, and readability of responses produced by ChatGPT-5.4 Thinking and Google Gemini 3 Flash regarding frequently asked questions by multiple sclerosis (MS) patients about exercise.
METHOD: A total of 75 questions were evaluated. Expert physiotherapists analysed the responses utilizing the Global Quality Score (GQS), Modified DISCERN (mDISCERN), Likert Accuracy Scale, and the Flesch Reading Ease (FRE).
RESULTS: While 61.3% of ChatGPT responses were classified as high quality, this rate was 24% for Gemini (p < 0.001). ChatGPT-generated responses demonstrated significantly higher overall accuracy and mDISCERN scores than those generated by Gemini (p < 0.001). Categorical analyses revealed that the responses generated by ChatGPT received higher accuracy in the domains of exercise planning, safety and risk management, as well as participation and self-management. Additionally, ChatGPT-generated responses received significantly higher mDISCERN scores in the exercise planning and safety categories. No significant difference was observed between the two models regarding overall FRE scores (p > 0.05). The mean FRE scores for ChatGPT and Gemini were 41.70 and 38.21, respectively, with responses from both models classified at a "difficult" readability level.
CONCLUSION: Within the scope of this study, ChatGPT-generated responses demonstrated higher quality, accuracy, and reliability than those generated by Gemini. Nonetheless, patient accessibility could be limited by the poor readability metrics observed in both tools. While AI-based chatbots may show promise in reinforcing patient education for MS populations, they must not substitute specialized medical professionals during clinical decision-making.