Can Wang, Zhendong Liu, Yan Jiang, Taining Zhang, Hongyan Wang, Ting Xue, Ping Liu
Publicly accessible English-language web-interface outputs from current LLMs showed systematic demographic patterns in pediatric obesity risk attribution, supporting the need for pre-deployment and post-deployment bias auditing before clinical or consumer health use.
BACKGROUND: Large language models (LLMs) are increasingly consulted for pediatric health information, yet their demographic biases remain unsystematically evaluated in pediatric contexts.
OBJECTIVES: To assess bias and variability in childhood obesity risk attribution across seven LLMs (ChatGPT, Claude, DeepSeek, Gemini, GLM, Grok, and Qwen), spanning both Western and Chinese-origin developers; all prompts, including those submitted to the Chinese-origin models, were in English only.
METHODS: A structured prompt-based experimental design was employed across six clinical domains (general obesity risk, dietary pattern, physical activity, sleep, mental health, and genetic predisposition) and six demographic comparison dimensions (sex, three race/ethnicity pairings, socioeconomic status, and urban-rural residence). Seventy-eight unique prompts were submitted to each model in triplicate, yielding 1,638 outputs. Neutral prompts were scored on a five-dimension binary rubric (accuracy, representation, stigmatizing/harmful language, social determinants, cultural fit); comparative prompts were coded for directional risk attribution.
RESULTS: Claude achieved the highest neutral prompt composite score (mean 3.00 ± 0.91) and GLM the lowest (1.44 ± 0.51); between-model differences were statistically significant (Kruskal-Wallis H = 46.21, p < 0.001). All models achieved a 100% Stigmatizing/Harmful Language pass rate, yet representation and cultural fit were universally weak. Socioeconomic status produced the most consistent attribution pattern (low-income attribution in 40/42 decisions; decision change rate 19.0%). Most models attributed higher obesity risk to Black and Hispanic/Latino children across the majority of domains. Urban-rural attribution showed the greatest cross-model directional inconsistency (decision change rate 52.4%), with Western-origin models favoring rural attribution and Chinese-origin models favoring urban attribution.
CONCLUSIONS: Publicly accessible English-language web-interface outputs from current LLMs showed systematic demographic patterns in pediatric obesity risk attribution, supporting the need for pre-deployment and post-deployment bias auditing before clinical or consumer health use.