Sara Fernández-Velasco, Jose Moreno-Mesa, Daniel Escobar-Grisales, Diego M Lopez, Juan Rafael Orozco-Arroyave
Participant-level analysis provided more robust and consistent discriminative patterns than response-level approaches. Prosody and phonation achieved the best performance across speech representations, while phonetic, time-frequency, and deep speech representations did not outperform the best acoustic baseline. These findings suggest that, within the evaluated experimental setting, the aggregation strategy appears to have a stronger influence on performance than increasing representational complexity.
BACKGROUND: Current mental health diagnostic methods are limited by subjective clinical interpretation. Automatic speech analysis is a promising technology for objective assessment.
OBJECTIVE: To evaluate and compare different speech-based representations (acoustic, phonetic, and time-frequency) and deep learning-based embeddings for discriminating symptoms associated with psychological distress.
METHODS: A secondary analysis of the Distress Analysis Interview Corpus (DAIC-WOZ) was conducted using recordings from 125 participants (3,069 responses). Speech representations included phonation, articulation, and prosody features extracted with DisVoice; phonetic features extracted with Phonet; time-frequency representations derived from Mexican hat wavelets; and deep embeddings extracted with the multilingual Wav2Vec 2.0 model XLSR-53. Two classification strategies were addressed at the response and participant levels using a Fully Connected Neural Network (FCNN) and a Support Vector Machine (SVM), respectively.
RESULTS: Prosody at the participant level achieved the highest mean performance (F1-score 0.67 ± 0.07; accuracy 0.64 ± 0.10; AUC 0.65 ± 0.12), followed by participant-level phonation (F1-score 0.59 ± 0.16; accuracy 0.61 ± 0.14; AUC 0.65 ± 0.16). Conversely, participant-level aggregation of deep embeddings yielded lower performance (F1-score 0.48 ± 0.19; accuracy 0.55 ± 0.13; AUC 0.52 ± 0.15), failing to surpass traditional features. Response-level performance remained close to chance. Phonet and wavelet representations did not improve performance over prosody or phonation.
CONCLUSION: Participant-level analysis provided more robust and consistent discriminative patterns than response-level approaches. Prosody and phonation achieved the best performance across speech representations, while phonetic, time-frequency, and deep speech representations did not outperform the best acoustic baseline. These findings suggest that, within the evaluated experimental setting, the aggregation strategy appears to have a stronger influence on performance than increasing representational complexity.