Yi-Lin Wang, Liao-Lin Chen, Ping-Ping Sun, Zhen-Zhen Ma, Ye-Ping Chen, Yang-Shuo Ge, Xin-Hui Huang, Chun-Meng Huang, Jia-Wei Du, Ting-Ting Meng, Dao-Fang Ding
Within this single-center, expert-rated benchmark, ChatGPT-o3 and DeepSeek performed comparably across most guideline-based Q&A dimensions, whereas ChatGPT-o3 achieved higher clinician-rated scores for draft rehabilitation-plan generation. These findings do not establish clinical effectiveness. Prospective patient-centered validation and clinician-supervised review are required before implementation in routine stroke rehabilitation.
OBJECTIVE: To evaluate large language models (LLMs) for stroke care in guideline-based question answering (Q&A) and individualized draft rehabilitation-plan generation, with explicit assessment of response quality, readability, and repeated-generation score stability.
METHODS: A two-stage study was conducted. Stage 1: Four LLMs generated best answers to 100 stroke-related questions; the top two models by accuracy advanced to Stage 2. In Stage 2a, ChatGPT-o3 and DeepSeek answered 20 guideline-derived clinical questions, each repeated three times. In Stage 2b, the two models generated individualized rehabilitation plans for 60 de-identified stroke cases, with three independent generations per case. Three senior clinicians evaluated outputs across correctness, completeness, readability, helpfulness, and safety using 5-point Likert scales. Chinese readability was assessed using the LDU-TGP platform, and English readability was assessed using the Flesch-Kincaid Grade Level. Generalized estimating equations were used for Stage 2a and Stage 2b comparisons to account for repeated generations clustered within clinical questions and patient cases, respectively.
RESULTS: In Stage 1, DeepSeek achieved the highest accuracy (91%), followed by ChatGPT-o3 (90%), Gemini (85%), and ChatGPT-4o (83%). In Stage 2a, no between-model differences across the five clinician-rated domains remained statistically significant after correction for multiple comparisons. ChatGPT-o3 generated clinical Q&A responses with a higher recommended reading age and English Flesch-Kincaid Grade Level, whereas the difference in Chinese Reading Difficulty Score did not remain significant after correction. In Stage 2b, ChatGPT-o3 significantly outperformed DeepSeek across correctness, completeness, readability, helpfulness, and safety in the GEE analysis. Repeated-generation score stability did not differ significantly between models.
CONCLUSIONS: Within this single-center, expert-rated benchmark, ChatGPT-o3 and DeepSeek performed comparably across most guideline-based Q&A dimensions, whereas ChatGPT-o3 achieved higher clinician-rated scores for draft rehabilitation-plan generation. These findings do not establish clinical effectiveness. Prospective patient-centered validation and clinician-supervised review are required before implementation in routine stroke rehabilitation.