José Ángel Martínez-Huertas, Alejandro Martínez-Mingo, Antonio López Rosell, Rodrigo S Kreitchmann, Guillermo Jorge-Botana
Surprisal is a measure from information theory that quantifies how unexpected an event is within the context of a particular system. In Large Language Models, higher surprisal of particular words indicates lower predictability or greater difficulty for cognitive processing. In language production, surprisal can be understood as a quantification of individual choices plus other produced linguistic characteristics such as selected content. Thus, this measure could be a quantifiable characteristic of how individuals produce their discourse and its predictability. However, it is not clear whether these measures can be considered an actual stable characteristic of the individuals or a measure of particular text answers without generalization to other language production of the same individual (i.e., a stable individual difference across language-based items). In a series of four studies, we analyzed the stability of surprisal measures across different language-based items of different tasks taken from previous publications. The reliability of mean surprisal scores was found to be excellent in the language-based task consisting of different language-based items, and to vary within the items themselves. We showed that surprisal is a stable individual characteristic regardless of the item but is significantly influenced by the length of the text answer. We also analyzed the relationship between mean surprisal scores and relevant variables ranging from language skills or domain-general cognitive skills to personality traits across the different datasets. We concluded that mean surprisal scores are related to certain cognitive functions and may also reflect alternative styles of speech or communication in speech production.