Junzhen Li, Yingjie Wu, Man Yang, Lihua Ren, Jian Qi, Hui Ouyang
current LLMs can provide useful general information about nutrition in cirrhosis, but none performed consistently well across safety, transparency, and readability. Their responses should supplement, not replace, guidance from hepatology and nutrition professionals.
BACKGROUND: large language models (LLMs) are increasingly used for public health information, but their performance in nutrition advice for cirrhosis remains uncertain. We investigated four LLMs in answering public questions about cirrhosis nutrition across safety, accuracy, empathy, information reliability and quality, and readability.
METHODS: in this cross-sectional comparative study, 50 public facing questions about nutrition in cirrhosis were submitted once to each of four LLMs (ChatGPT 5.5, Claude Opus 4.7, DeepSeek V4 Pro, and Gemini 3.1 Pro Thinking) between May 10 and May 16, 2026. Three raters independently assessed 200 responses for safety, accuracy, empathy, and information quality. Readability was assessed with six indices.
RESULTS: a total of 200 responses were evaluated. Safety coding showed high inter-rater agreement (Fleiss' kappa = 0.864; 95 % CI, 0.755-0.958). ICC values for other manually scored outcomes ranged from 0.860 to 0.892. Potentially unsafe responses occurred in all models, ranging from 6.0 % to 12.0 %, with no significant difference across models (raw p = 0.599). Accuracy differed significantly across models (raw p = 0.003), with the highest median score for ChatGPT 5.5. Empathy also differed significantly (raw p < 0.001), with higher median scores for Claude Opus 4.7 and Gemini 3.1 Pro Thinking. DISCERN, EQIP, JAMA and all readability indices differed significantly across models, while GQS did not.
CONCLUSION: current LLMs can provide useful general information about nutrition in cirrhosis, but none performed consistently well across safety, transparency, and readability. Their responses should supplement, not replace, guidance from hepatology and nutrition professionals.