科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ The Knee2026-09-11

Comparative quality, accuracy, and readability of large language model responses to patient questions about robotic-assisted total knee arthroplasty.

Ulas Can Kolac, Mazlum Veysel Sili, Orhan Mete Karademir, Gokhan Ayik, Mustafa Akkaya, Gökhan Çakmak

一句话结论 · In one sentence

The evaluated LLMs demonstrated generally acceptable clinical accuracy but differed across measures of written information quality, understandability, and readability. No significant difference was detected using QAMAI. Although Gemini 3 and DeepSeek demonstrated greater readability in selected comparisons, median responses across all models remained above recommended patient-education reading levels. LLM-generated responses should therefore be regarded as supplementary rather than standalone sources of patient information regarding RA-TKA.

原始摘要(英文原文)· Original abstract
PURPOSE: To compare the information quality, accuracy, and readability of patient-directed responses generated by large language models (LLMs), including ChatGPT-o3, ChatGPT-5.2, Gemini 3, and DeepSeek, regarding robotic-assisted total knee arthroplasty (RA-TKA). METHODS: Thirty frequently asked patient questions were identified using LLM outputs and Google search queries. Responses were evaluated for information quality using the DISCERN and Quality Analysis of Medical Artificial Intelligence (QAMAI) instruments, for clinical accuracy using a 5-point ordinal rating scale, and for understandability and readability using the PEMAT Understandability and Flesch-Kincaid Reading Ease scores. RESULTS: Median DISCERN scores were 46.0 (range, 35.0-50.0) for ChatGPT-o3, 45.75 (28.5-51.0) for ChatGPT-5.2, 43.75 (32.0-47.5) for Gemini 3, and 42.0 (32.0-50.0) for DeepSeek, with a significant overall difference among models (p < 0.001). The 5-point clinical accuracy scores were similar across models (median 4.0; p = 0.636). Median QAMAI scores were 23.0 for all four models, without a significant between-model difference (p = 0.462). PEMAT Understandability scores differed significantly among models (p < 0.001), with median scores of 90.0 for ChatGPT-o3 and ChatGPT-5.2, 88.0 for Gemini 3, and 85.0 for DeepSeek. Flesch-Kincaid Reading Ease scores also differed significantly (p < 0.001); Gemini 3 demonstrated higher readability than both ChatGPT models, whereas DeepSeek demonstrated higher readability than ChatGPT-o3. CONCLUSION: The evaluated LLMs demonstrated generally acceptable clinical accuracy but differed across measures of written information quality, understandability, and readability. No significant difference was detected using QAMAI. Although Gemini 3 and DeepSeek demonstrated greater readability in selected comparisons, median responses across all models remained above recommended patient-education reading levels. LLM-generated responses should therefore be regarded as supplementary rather than standalone sources of patient information regarding RA-TKA.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Comparative quality, accuracy, and readability of large language model responses to patient questions about robotic-assisted total knee arthroplasty. — 科研速览 Science Skim