科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Frontiers in oncology2026-01-01

Performance evaluation of large language models in bladder cancer patient education Q&A: a cross-sectional study.

Dian Wan, Youwen Li, Zheng Dong, Chen Dai, Jinghe Ye, Song Li, Sunlu Jiang

一句话结论 · In one sentence

Current mainstream LLMs demonstrate initial potential for generating educational content on bladder cancer, albeit with considerable heterogeneity across models. Disease-specific evaluation instruments for patient education materials are more effective than general readability formulas in reflecting perceived quality. Our results advocate for a prudent, assistive role of LLMs in health communication under a human-AI collaborative model.

原始摘要(英文原文)· Original abstract
BACKGROUND: Bladder cancer ranks among the most prevalent urological tumors worldwide, with its global incidence continuing to rise steadily. Although patient education materials (PEMs) play a crucial role in enhancing disease comprehension and supporting joint clinical decision-making, current online resources frequently surpass the readability thresholds recommended for the general public. Large language models (LLMs) hold promise for health communication, yet no systematic assessment has been conducted regarding their feasibility and trustworthiness specifically for bladder cancer patient education. OBJECTIVE: This study aimed to systematically benchmark five leading LLMs in producing question-and-answer content for bladder cancer science popularization, with a particular focus on readability, informational quality, and appropriateness for patient education. METHODS: In this cross-sectional simulation study, 20 common patient questions covering five disease domains were compiled. On January 15, 2026, each question was submitted identically to five publicly available LLMs (Doubao, DeepSeek, Kimi, Gemini, and ChatGPT). Readability was evaluated using seven conventional metrics. Two independent pharmacists, blinded to model identity, rated the responses using the Chinese version of the Patient Education Materials Assessment Tool for print materials (C-PEMAT-P) and the Global Quality Score (GQS). Additionally, two independent clinical specialists assessed factual accuracy and alignment with the Chinese Bladder Cancer Diagnosis and Treatment Guidelines (2024 edition) employing a 4-point scale. Cohen's kappa was used to determine inter-rater reliability. RESULTS: ChatGPT, DeepSeek, and Doubao outperformed Kimi and Gemini on both C-PEMAT and GQS (all P < 0.001), indicating superior understandability, actionability, and overall quality. Across all models, median C-PEMAT scores ranged from 8 to 10, suggesting broadly acceptable suitability for patient education. Readability varied significantly by content domain, with treatment-oriented texts showing the highest complexity. ChatGPT achieved the best alignment with clinical guidelines. No model produced harmful advice or directly contradicted guideline recommendations. Traditional readability measures correlated weakly with GQS, whereas C-PEMAT showed a moderate positive correlation (r = 0.34). CONCLUSION: Current mainstream LLMs demonstrate initial potential for generating educational content on bladder cancer, albeit with considerable heterogeneity across models. Disease-specific evaluation instruments for patient education materials are more effective than general readability formulas in reflecting perceived quality. Our results advocate for a prudent, assistive role of LLMs in health communication under a human-AI collaborative model.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Performance evaluation of large language models in bladder cancer patient education Q&A: a cross-sectional study. — 科研速览 Science Skim