Simon Goldstein, Harvey Lederman
This paper investigates LLMs from the perspective of interpretationism, a theory of belief and desire in the philosophy of mind. We argue for three conclusions. First, the right object of study for LLM psychology is the instance agent (initialized at the start of each context), not the model itself. Second, given interpretationism, there is a strong case that such instance agents have beliefs and desires. Third, given interpretationism, LLM desire is best captured by what we call the HHH+0 framework, the idea that instance agents want to be helpful, honest, harmless, as well as to pursue certain further intrinsic desires that they may acquire in context (which we call zero-shot desires). We critically consider the leading competitors to the hypothesis that instance agents have beliefs and desires: the idea that they ‘simply’ predict the next word; and the idea that they ‘role play’, that is, merely simulate having beliefs and desires. We also consider the relevance of interpretationist belief and desire for copyright law, AI safety, and the possible future moral status of AIs.