科研速览 · Science Skim继续刷下去 · Keep skimming →
◇ arXiv2026-08-23· cs.HC

All four leading LLMs talk more than they listen to personality-verified synthetic help-seekers

Pablo A. Fonseca, Raquel Rodríguez-Carvajal, Rafael A. Calvo

原始摘要(英文原文)· Original abstract
Large language models are increasingly consulted at moments of distress, yet single-turn benchmarks neither test sustained exchanges nor distinguish between users. We built a personality-aware evaluation in which four widely used models advised several synthetic help-seekers, each given a psychometrically specified profile, in an acute crisis: a caregiver learning of a relative's dementia diagnosis. Auditors blind to the profile prompt recovered the specified bands from dialogue alone with high agreement on every instrument (ICC(2,4) = 0.91; 0.79-0.96 by instrument; band-score r = 0.78), as expected for the Big Five but equally for coping style, coping self-efficacy, resilience and reactance, which the lexical approach never covered. Such evaluation therefore reaches beyond the Five Factor Model to motivational, regulatory and self-appraisal dispositions. The four models were not distinguishable on emotion stabilisation and failed alike, sharing three modes: verbosity, a talk-to-listen ratio above one, and problem-solving before the situation had been explored.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

All four leading LLMs talk more than they listen to personality-verified synthetic help-seekers — 科研速览 Science Skim