科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Urology2026-08-17

Evaluation of AI Chatbot Responses to Pediatric Urology Frequently Asked Questions.

Najva Mazhari, Andrew Freedman, Nadine Friedrich, Timothy Daskivich, Paul Kokorowski

一句话结论 · In one sentence

LLMs can provide accurate, safe, and supportive information for parents, but their usefulness is limited by gaps in completeness and high readability levels. Claude achieved the highest overall quality, whereas ChatGPT demonstrated high safety with lower completeness. Future development would prioritize plain-language optimization, context-aware emotional framing, and parent co-design to improve comprehension.

原始摘要(英文原文)· Original abstract
OBJECTIVE: To evaluate the quality of responses from four publicly available LLMs (ChatGPT-4o, Claude 3.7, Gemini 2.5, and Copilot) to frequently asked questions (FAQs) in pediatric urology. METHODS: FAQs were generated using standardized prompts and submitted to each LLM using parent-centered instructions. Two board-certified pediatric urologists independently rated responses across seven domains: Accuracy, Completeness, Safety, Clarity, Actionability, Conciseness, and Global Quality Score using a 5-point Likert scale. Emotional tone was analyzed using the NRC Word-Emotion Association Lexicon, and readability was assessed using seven established metrics. RESULTS: All LLMs produced generally high expert ratings for accuracy (mean 4.27), safety (4.36), and clarity (4.42). Claude achieved the highest overall quality score (4.35), followed by Gemini (4.03), ChatGPT (4.00), and Copilot (3.80). Emotional tone was predominantly positive and supportive across models. Reading corresponded to an 11th-15th grade reading level, exceeding the recommended 6th-8th grade patient-education standards. CONCLUSION: LLMs can provide accurate, safe, and supportive information for parents, but their usefulness is limited by gaps in completeness and high readability levels. Claude achieved the highest overall quality, whereas ChatGPT demonstrated high safety with lower completeness. Future development would prioritize plain-language optimization, context-aware emotional framing, and parent co-design to improve comprehension.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Evaluation of AI Chatbot Responses to Pediatric Urology Frequently Asked Questions. — 科研速览 Science Skim