科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Nutricion hospitalaria2026-09-16

Performance of large language models in answering public questions about nutrition in cirrhosis: a comparative study.

Junzhen Li, Yingjie Wu, Man Yang, Lihua Ren, Jian Qi, Hui Ouyang

一句话结论 · In one sentence

current LLMs can provide useful general information about nutrition in cirrhosis, but none performed consistently well across safety, transparency, and readability. Their responses should supplement, not replace, guidance from hepatology and nutrition professionals.

原始摘要(英文原文)· Original abstract
BACKGROUND: large language models (LLMs) are increasingly used for public health information, but their performance in nutrition advice for cirrhosis remains uncertain. We investigated four LLMs in answering public questions about cirrhosis nutrition across safety, accuracy, empathy, information reliability and quality, and readability. METHODS: in this cross-sectional comparative study, 50 public facing questions about nutrition in cirrhosis were submitted once to each of four LLMs (ChatGPT 5.5, Claude Opus 4.7, DeepSeek V4 Pro, and Gemini 3.1 Pro Thinking) between May 10 and May 16, 2026. Three raters independently assessed 200 responses for safety, accuracy, empathy, and information quality. Readability was assessed with six indices. RESULTS: a total of 200 responses were evaluated. Safety coding showed high inter-rater agreement (Fleiss' kappa = 0.864; 95 % CI, 0.755-0.958). ICC values for other manually scored outcomes ranged from 0.860 to 0.892. Potentially unsafe responses occurred in all models, ranging from 6.0 % to 12.0 %, with no significant difference across models (raw p = 0.599). Accuracy differed significantly across models (raw p = 0.003), with the highest median score for ChatGPT 5.5. Empathy also differed significantly (raw p < 0.001), with higher median scores for Claude Opus 4.7 and Gemini 3.1 Pro Thinking. DISCERN, EQIP, JAMA and all readability indices differed significantly across models, while GQS did not. CONCLUSION: current LLMs can provide useful general information about nutrition in cirrhosis, but none performed consistently well across safety, transparency, and readability. Their responses should supplement, not replace, guidance from hepatology and nutrition professionals.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Performance of large language models in answering public questions about nutrition in cirrhosis: a comparative study. — 科研速览 Science Skim