Yanru Jiang, Qianyun Wang, Liang Zheng, Lei Zhang, Bo Jiang, Jingjuan Xu, Jun Wang, Jiaying Zhu, Zhishui Wu
Public-facing chatbots differed substantially in safety, reliability, communication quality, and readability. These findings are time-, interface-, and prompt-dependent. Chatbots may support general patient education but should not replace individualized clinician-led prognostic communication.
OBJECTIVE: To compare the safety, accuracy, empathy, reliability, information quality, and readability of five publicly accessible large language model chatbots when answering patient-facing lung cancer prognostic questions under standardized single-turn English prompting.
METHODS: In this Chatbot Health Advice Reporting Transparency-guided cross-sectional evaluation, 53 standardized English prompts were submitted once to ChatGPT, Gemini, Copilot, DeepSeek, and Doubao through official web interfaces during April 1-21, 2026. Five blinded raters assessed 265 responses for safety, accuracy, empathy, DISCERN, EQIP, JAMA benchmark criteria, Global Quality Scale, and readability. Paired repeated-measures analyses were used.
RESULTS: Inter-rater agreement was good to excellent. Safety differed significantly across models (Cochran's Q = 14.089, df = 4, p = 0.007). Gemini generated the highest proportion of safe responses (48/53, 90.6%), whereas DeepSeek generated the lowest (33/53, 62.3%). The only adjusted pairwise safety difference that remained significant was Gemini versus DeepSeek (adjusted p = 0.023). Accuracy, empathy, reliability, information quality, and readability also differed significantly across models (all p < 0.001). Gemini showed the most favorable descriptive profile for safety, accuracy, empathy, and reliability, while Copilot produced the most readable responses.
CONCLUSION: Public-facing chatbots differed substantially in safety, reliability, communication quality, and readability. These findings are time-, interface-, and prompt-dependent. Chatbots may support general patient education but should not replace individualized clinician-led prognostic communication.