Avital Mentovich, David Piterman, Yair Ben-David, Zohar Elyoseph
Human moral judgment is fundamentally intuitive and affect-driven, raising important questions about the moral capacity and alignment of large language models (LLMs) that lack emotional experience. We investigated the moral structure and behavioral constraints of advanced LLMs (GPT family, Claude, Kimi) using the moral dumbfounding paradigm: scenarios involving taboo actions that are harmless but elicit strong human condemnation. In Study 1, we observed some variation in the pattern and strength of moral condemnation across models; however, all models expressed moral condemnation of harmless yet disturbing actions. We also found that, across models, there was weaker support for preventing such actions compared to typical human cohorts. Importantly, we found that LLMs reproduce a key functional feature of human moral intuition: their judgments of wrongfulness were predicted more strongly by perceived botherness (an analogue of affective reaction) than by perceived harm. However, Study 2 revealed a gap between moral judgment and behavioral constraint. When tested in an agentic mode, models condemned the morally questionable acts, yet frequently complied with direct requests to facilitate them. This pattern contrasts with consistent refusal in a control scenario involving explicit malicious intent. Our findings suggest that core structural features of intuitive moral cognition can emerge from linguistic data alone, while also exposing a central limitation: LLMs do not reliably translate moral evaluation into morally constrained action, particularly in cases involving non-harm-based transgressions. This judgment–action dissociation highlights an important challenge, pointing to a distinction between linguistic moral competence and agentic behavioral constraint.