科研速览继续刷下去 →
◆ Information Processing & Management2026-09-26· Psychology

PhDGPT: Measuring bias in academic mental health across quantised and uncensored LLMs

Edoardo Sebastiano De Duro, Enrique Taietta, Simon D’Alfonso, Arianna Bentenuto, Giulio Rossetti, Massimo Stella

一句话结论

We investigate bias in academic mental health as reflected by Large Language Models (LLMs).

原始摘要(原文)
We investigate bias in academic mental health as reflected by Large Language Models (LLMs). To this aim, we introduce PhDGPT, a synthetic dataset encapsulating the machine psychology of Ph.D. researchers and professors as perceived by OpenAI’s GPT-3.5 and GPT-OSS-20b. The dataset consists of 18,000 data points for each model, comprising 300 iterations repeated across 15 academic events, 2 gender personas, 2 career levels and 42 unique items responses of the Depression, Anxiety, and Stress Scale (DASS-42) across 2 LLMs. The PhDGPT dataset integrates these psychometric scores with their explanations in plain language. This synergy of scores and texts offers a dual, comprehensive perspective on the emotional well-being of simulated academics, e.g. male/female Ph.D. students or professors. We find that both GPT-3.5 and GPT-OSS produce higher levels of mental distress when impersonating Ph.D. students compared to professors, a direction that aligns with humans’ recent results about mental health in academia. Compared against more than 39,000 human psychometric responses, GPT-3.5 reproduces the canonical structure of anxiety, stress and depression, whereas GPT-OSS does so only under vLLM inference and not under the GGUF quantisation analysed here. This specific instance of GPT-OSS, run via LM Studio under GGUF quantisation, also produces markedly fewer linguistic patterns correlating with psychological dimensions of distress than GPT-3.5. A robustness check — encompassing prompt paraphrasing across four alternative formulations and 3 additional open-source models spanning a range of alignment strategies (Mistral Small 24b Instruct, and 2 uncensored variants) — confirms that psycholinguistic signatures are stable across surface-level prompt variation, while revealing that the relationship between safety alignment and psycholinguistic richness is more complex than a simple censorship effect. We treat safety alignment as one of several candidate explanations for the observed degradation, alongside architectural differences and training procedures.
读原文 ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文

PhDGPT: Measuring bias in academic mental health across quantised and uncensored LLMs — 科研速览 Science Skim