Jun Ji, Yantao Ma
"VoiceDiaryMood is a restricted-access benchmark of 450 de-identified Chinese transcripts of spoken emotion diaries recorded by 37 patients with mood disorders in an outpatient follow-up program (2018\u20132022). Each transcript was independently annotated by three psychiatrists using a 12-field ordinal schema covering suicide-risk tier, suicidality, and ten symptom dimensions anchored to the C-SSRS, PHQ-9, GAD-7, and YMRS, including rarely benchmarked mania-spectrum dimensions. The release provides per-rater labels with explicit \"not assessable\" coding, aggregated gold labels (19 high-risk diaries), patient-grouped five-fold cross-validation assignments, per-diary outputs of four zero-shot large language model arms and a fine-tuned local model, and a curated catalogue of automatic-speech-recognition misrecognitions, including homophone errors that mask suicidal ideation. Transcripts are de-identified through a reviewed two-pass pipeline; platform identifiers are replaced by random IDs; raw audio and the diary-patient mapping are never distributed. Access requires a data use agreement: research use only, no re-identification, no redistribution, and no use in patient-facing clinical decision products. The dataset supports research on clinician-aligned risk screening, annotation disagreement, and privacy-preserving clinical NLP."