科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Breastfeeding medicine : the official journal of the Academy of Breastfeeding Medicine2026-08-20

Accuracy and Safety of Language Model-Generated Breastfeeding Counseling Responses: An Expert-Based Comparative Evaluation of ChatGPT-3.5, ChatGPT-4, and BreastfeedGPT.

Seda Serhatlıoğlu, Yeşim Yeşil

一句话结论 · In one sentence

BreastfeedGPT may offer advantages over general-purpose AI tools in providing breastfeeding-related information. However, because human expert responses were not evaluated in parallel, its role as a complementary resource should be interpreted cautiously and confirmed in future direct comparisons. Such tools should support, rather than replace, professional breastfeeding care.

原始摘要(英文原文)· Original abstract
AIM: This study aims to compare the responses generated by ChatGPT-3.5, ChatGPT-4, and BreastfeedGPT, a customized domain-specific GPT, to frequently asked questions related to breastfeeding counseling. DESIGN AND METHODS: Ten breastfeeding-related questions were selected based on international guidelines and expert validation. Each model generated responses to the same questions, which were anonymized and evaluated by eight health professionals with expertise in breastfeeding counseling using structured Likert scales. Evaluation dimensions included (1) scientific accuracy and scope, (2) tone, language, and motivation, and (3) reliability via mDISCERN scoring. Friedman and Wilcoxon signed-rank tests with Bonferroni correction were used to determine statistical significance. RESULTS: BreastfeedGPT model outperformed ChatGPT-3.5 and ChatGPT-4 across all domains. It received the highest mean scores in scientific accuracy (3.72 ± 0.23) and motivational tone (3.78 ± 0.42), with statistically significant differences (p < 0.001). Although the mDISCERN scores also favored the custom model, pairwise differences were not statistically significant after correction. GPT-4 performed better than GPT-3.5, but worse than BreastfeedGPT. CONCLUSION: BreastfeedGPT may offer advantages over general-purpose AI tools in providing breastfeeding-related information. However, because human expert responses were not evaluated in parallel, its role as a complementary resource should be interpreted cautiously and confirmed in future direct comparisons. Such tools should support, rather than replace, professional breastfeeding care.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Accuracy and Safety of Language Model-Generated Breastfeeding Counseling Responses: An Expert-Based Comparative Evaluation of ChatGPT-3.5, ChatGPT-4, and BreastfeedGPT. — 科研速览 Science Skim