Anling Xiang
Large language models (LLMs) have achieved remarkable performance gains, yet concerns about their embedded biases remain pressing. Existing research often targets a single dimension, lacking systematic comparisons across different types of bias. This study introduces a three-level diagnostic framework—explicit, evaluative, and implicit—to characterize the “bias iceberg” in LLMs. We construct a bilingual dataset (Chinese–English) spanning seven socio-psychological dimensions (appearance, competence, dominance, emotion, leadership, morality, and physicality), comprising approximately 420 minimal-pair sentences and 400 word-association sets, and conduct unified evaluations across eight mainstream models (GPT-4o, Claude-3.7, Gemini-2.5, Grok-3, Qwen-Plus, DeepSeek-v3, Doubao, and Kimi-2). Results reveal that explicit bias remains generally low ( mean ≈ 0.0141), though residual disadvantages for women persist in appearance and emotion. Evaluative bias intensifies to a moderate level ( mean ≈ 0.0259), with directional divergence in morality and dominance. Implicit bias emerges as the most pronounced ( mean ≈ 0.1738, peaking at 0.45), manifesting stable male-anchoring effects in dominance and physicality, with a magnitude 12.3 times greater than explicit bias. Network analysis uncovers three structural archetypes—hyper-centralized, balanced clustering, and de-centralized. Cross-cultural comparisons further show that U.S. models more strongly reproduce “male–power/physicality” associations at evaluative and implicit levels, whereas Chinese models exhibit greater convergence. The proposed framework and dataset are reproducible and cross-culturally adaptable, offering new empirical evidence and structured insights for uncovering and mitigating deep-seated biases in LLMs.