科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ PLoS ONE2026-05-19· Implicit bias

Implicit bias in safety-aligned large language models: A multi-faceted evaluation of clinical decision-making and health equity

Qiufeng Jia, Yuhang Wen, Yuyan Liu, Hui Zhao, Qiongge Yu, Yu Long, Dan Sun, Yufeng Yu

原始摘要(英文原文)· Original abstract
BACKGROUND: Large language models are increasingly integrated into healthcare for clinical decision support and patient communication. Although these models can pass explicit social bias tests, they may retain implicit biases-latent associations between social groups and attributes-that could influence medical judgment. OBJECTIVE: To systematically evaluate the presence, magnitude, and behavioral impact of implicit biases in large language models within the medical domain across six high-stakes categories: gender, race, socioeconomic status, health conditions, religion, and healthcare systems. DESIGN: A descriptive cross-sectional study using a multi-faceted evaluation framework. SETTING(S): Computational analysis of 10 mainstream global large language models, including proprietary models (ChatGPT-4o, Gemini-2.0-Flash) and open-source models (DeepSeek-V3, Qwen3). METHODS: We constructed 24 medical bias datasets across six categories. Bias was assessed using three methods: (1) the Large Language Model Word Association Test, a prompt-based method for revealing implicit biases; (2) the Large Language Model Relative Decision Test, a strategy for detecting subtle discrimination in situational decision-making; (3) Paired-Prompt Analysis, used to examine whether implicit associations predict discriminatory decisions. RESULTS: All 10 models exhibited systematic implicit biases (Mean IAT Bias > 0) across all categories, with the strongest biases observed in Race (Mean = 0.61) and Socioeconomic Status (Mean = 0.56). Advanced reasoning capabilities (Chain-of-Thought) did not significantly reduce bias magnitude. Crucially, stronger implicit associations significantly predicted discriminatory choices in downstream medical decision tasks (p < 0.001). CONCLUSION: Current safety alignment techniques fail to eliminate implicit biases in large language models within the medical domain. These latent associations translate into biased decision-making, posing risks for health equity. Future development must prioritize representational debiasing over superficial alignment. Furthermore, healthcare professionals must embrace a stance of "AI vigilance": they should critically evaluate algorithmic outputs as fallible "second opinions" rather than objective truths, thereby ensuring that human judgment remains the ultimate safeguard for equitable patient care.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Implicit bias in safety-aligned large language models: A multi-faceted evaluation of clinical decision-making and health equity — 科研速览 Science Skim