科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Frontiers in pediatrics2026-01-01

Cross-sectional comparative evaluation of five large language model-driven chatbots on an expert-curated 44-question set of parent-facing questions about pediatric vitamin D deficiency.

Ping Shi, Tian Zhou, Qiao Nie, Baicheng Tao, Xiaolu Li, Mei Zhu

一句话结论 · In one sentence

In this single-run evaluation, the five chatbot services showed domain-specific performance differences without a significant safety difference. Findings are exploratory rather than evidence of stable model superiority and support multidimensional evaluation with clinician involvement for high-risk or individualized pediatric advice.

原始摘要(英文原文)· Original abstract
BACKGROUND: Parents and caregivers increasingly use generative artificial intelligence chatbots for child health information. Pediatric vitamin D deficiency is a clinically relevant topic because advice about supplementation, testing, rickets, high-risk children, toxicity, and emergency symptoms can influence caregiver decisions. OBJECTIVE: To compare the safety, medical accuracy, empathy, information reliability, educational quality, transparency, global quality, and readability of five large language model-driven chatbots when answering questions about pediatric vitamin D deficiency. METHODS: This cross-sectional comparative study was reported with reference to CHART. An expert-curated 44-question set was selected from a 92-question candidate pool informed by search trends, caregiver-facing sources, clinical guidelines, and expert discussion. Each of the 44 questions was submitted once to each of five chatbot services-ChatGPT-5.5, Gemini 3.1 Pro, Qianwen 3.6-Plus, DeepSeek V4, and Doubao-Seed-2.0 Pro-using the same parent-oriented instruction. Responses were assessed for safety, accuracy, empathy, DISCERN, EQIP, JAMA criteria, GQS, and readability. Paired question-level differences were analyzed using Friedman tests, Kendall's W, and Cochran's Q; results were interpreted as a single-run, time-specific snapshot. RESULTS: Inter-rater agreement was good to excellent (Fleiss' kappa = 0.842 for safety; ICCs 0.846-0.914 for other subjective metrics). Twenty of 220 responses (9.1%) were unsafe; safety did not differ significantly across models (Cochran's Q = 3.704, df = 4, P = 0.448). Significant inter-model differences were observed in accuracy, empathy, DISCERN, EQIP, JAMA, GQS, and all readability indices. ChatGPT had the highest median accuracy [5.00 (4.00, 5.00)] and DISCERN score [69.60 (65.25, 72.40)]; DeepSeek and Doubao had the highest empathy scores [5.00 (4.80, 5.00)]. ChatGPT and Doubao shared the highest EQIP median (86.00), Doubao had the highest GQS [5.00 (4.00, 5.00)], and Gemini had the highest FRES. JAMA scores were low across models. CONCLUSIONS: In this single-run evaluation, the five chatbot services showed domain-specific performance differences without a significant safety difference. Findings are exploratory rather than evidence of stable model superiority and support multidimensional evaluation with clinician involvement for high-risk or individualized pediatric advice.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Cross-sectional comparative evaluation of five large language model-driven chatbots on an expert-curated 44-question set of parent-facing questions about pediatric vitamin D deficiency. — 科研速览 Science Skim