科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ PNAS nexus2026-08-01

Like humans, language models demonstrate face-to-character biases.

Steven A Lehr, Yash Lothe, Mahzarin R Banaji

原始摘要(英文原文)· Original abstract
Humans routinely make unjustified character inferences, such as labeling people as trustworthy or untrustworthy, based on facial features. We conducted 13 experiments, using four models and totaling nearly 8,000 trials, to ask: Would large language models (LLMs), trained foundationally on language and not images, mirror human biases by inferring character traits from 2D facial images, or would they be free from this human error? GPT-4o reliably exhibited face-biased judgments of competence and trustworthiness (experiments 1 and 2), generalized these judgments to semantically related traits (experiments 3 and 4), and even to primate faces that were absent from its training data (experiment 5). Strikingly, the LLM's face-to-character inferences escalated to ascriptions of extreme negative behavior such as murder and human trafficking (experiment 6) as well as to positive high-impact decisions like selection for employment or venture capital funding (experiment 7). To ensure these results were not an idiosyncratic feature of a particular LLM (GPT-4o), in experiments 8-13, we showed the same patterns, usually with an even greater degree of bias, in GPT-5, Gemini 3 Flash Preview, and Claude Sonnet 4.5. These robust biases stand in contrast to LLMs' reluctance to explicitly endorse race/gender stereotypes, and suggest that alignment efforts to date have been domain-specific, with current models lacking generalized egalitarian decision-making.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Like humans, language models demonstrate face-to-character biases. — 科研速览 Science Skim