科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Frontiers in digital health2026-01-01

Automatic analysis of speech representations to assess psychological distress.

Sara Fernández-Velasco, Jose Moreno-Mesa, Daniel Escobar-Grisales, Diego M Lopez, Juan Rafael Orozco-Arroyave

一句话结论 · In one sentence

Participant-level analysis provided more robust and consistent discriminative patterns than response-level approaches. Prosody and phonation achieved the best performance across speech representations, while phonetic, time-frequency, and deep speech representations did not outperform the best acoustic baseline. These findings suggest that, within the evaluated experimental setting, the aggregation strategy appears to have a stronger influence on performance than increasing representational complexity.

原始摘要(英文原文)· Original abstract
BACKGROUND: Current mental health diagnostic methods are limited by subjective clinical interpretation. Automatic speech analysis is a promising technology for objective assessment. OBJECTIVE: To evaluate and compare different speech-based representations (acoustic, phonetic, and time-frequency) and deep learning-based embeddings for discriminating symptoms associated with psychological distress. METHODS: A secondary analysis of the Distress Analysis Interview Corpus (DAIC-WOZ) was conducted using recordings from 125 participants (3,069 responses). Speech representations included phonation, articulation, and prosody features extracted with DisVoice; phonetic features extracted with Phonet; time-frequency representations derived from Mexican hat wavelets; and deep embeddings extracted with the multilingual Wav2Vec 2.0 model XLSR-53. Two classification strategies were addressed at the response and participant levels using a Fully Connected Neural Network (FCNN) and a Support Vector Machine (SVM), respectively. RESULTS: Prosody at the participant level achieved the highest mean performance (F1-score 0.67 ± 0.07; accuracy 0.64 ± 0.10; AUC 0.65 ± 0.12), followed by participant-level phonation (F1-score 0.59 ± 0.16; accuracy 0.61 ± 0.14; AUC 0.65 ± 0.16). Conversely, participant-level aggregation of deep embeddings yielded lower performance (F1-score 0.48 ± 0.19; accuracy 0.55 ± 0.13; AUC 0.52 ± 0.15), failing to surpass traditional features. Response-level performance remained close to chance. Phonet and wavelet representations did not improve performance over prosody or phonation. CONCLUSION: Participant-level analysis provided more robust and consistent discriminative patterns than response-level approaches. Prosody and phonation achieved the best performance across speech representations, while phonetic, time-frequency, and deep speech representations did not outperform the best acoustic baseline. These findings suggest that, within the evaluated experimental setting, the aggregation strategy appears to have a stronger influence on performance than increasing representational complexity.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Automatic analysis of speech representations to assess psychological distress. — 科研速览 Science Skim