科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Frontiers in public health2026-01-01

Evaluating search-enabled large language model interfaces for mpox public health consultation: a guideline-based comparative study.

Qiqi Zheng, Ru Chen, Mingming Cai, Yanyan Chen, Xiuli Lin, Tingting Wang

一句话结论 · In one sentence

The evaluated search-enabled LLM interfaces showed heterogeneous performance across safety, accuracy, empathy, reliability/information quality, and readability. Although unsafe responses were relatively uncommon, potentially harmful outputs occurred in every interface. These findings support the need for guideline-based evaluation, source transparency, readability optimization, and robust safety safeguards when such interfaces are evaluated or considered for mpox-related public health consultation. The results represent a time- and configuration-specific interface-level snapshot; they should not be attributed to the underlying base models in isolation or interpreted as establishing reproducible performance or a stable hierarchy across sessions, versions, or settings.

原始摘要(英文原文)· Original abstract
BACKGROUND: Search-enabled large language model interfaces are increasingly used by the public for health information, but their performance in mpox-related public health consultation remains unclear. This study evaluated their safety, accuracy, empathy, reliability/information quality, and readability. METHODS: We conducted a single-query comparative cross-sectional evaluation using 52 predefined mpox-related public consultation questions. Each question was submitted once to each of six search-enabled LLM interfaces, yielding 312 first responses. Responses were assessed against a guideline-based reference framework. Safety was coded as a binary outcome, while accuracy and empathy were rated on 5-point scales. Reliability/information quality was evaluated using DISCERN, EQIP, JAMA benchmark criteria, and GQS. Readability was assessed using six established readability indices. Five trained raters independently evaluated the human-scored outcomes. RESULTS: Unsafe responses were relatively infrequent but occurred in all six interfaces, with safe-response rates ranging from 86.5 to 92.3%. No pairwise difference in Safety remained statistically significant after Benjamini-Hochberg correction. Overall differences across interfaces were statistically significant for Accuracy, Empathy, all four reliability/information quality measures, and all six readability indices. Benjamini-Hochberg-adjusted post hoc analyses identified outcome-specific pairwise differences, although the pairwise patterns varied across measures. CONCLUSION: The evaluated search-enabled LLM interfaces showed heterogeneous performance across safety, accuracy, empathy, reliability/information quality, and readability. Although unsafe responses were relatively uncommon, potentially harmful outputs occurred in every interface. These findings support the need for guideline-based evaluation, source transparency, readability optimization, and robust safety safeguards when such interfaces are evaluated or considered for mpox-related public health consultation. The results represent a time- and configuration-specific interface-level snapshot; they should not be attributed to the underlying base models in isolation or interpreted as establishing reproducible performance or a stable hierarchy across sessions, versions, or settings.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Evaluating search-enabled large language model interfaces for mpox public health consultation: a guideline-based comparative study. — 科研速览 Science Skim