科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Frontiers in public health2026-01-01

A multimodal feature fusion model integrating voice acoustic features and early key risk factors for early screening of postpartum depression: a prospective short-term longitudinal observational study.

Xuefei Han, Chongyu Yue, Xinwei Zhang, Xiaofei Ji, Shusen Lin, Meiyu Chen, Xiaojing Wang, Yanxia Zhang, Wanyu Xu, Yinhua Liu, Huawei Li

一句话结论 · In one sentence

Voice acoustic features collected before hospital discharge show potential as objective markers for early PPD screening, although further validation in larger, independent cohorts is required before clinical application.

原始摘要(英文原文)· Original abstract
BACKGROUND: Postpartum depression (PPD) is a prevalent perinatal mental health disorder, yet early identification remains challenging due to the reliance on subjective self-report screening. Voice acoustic features offer a promising objective alternative, but their utility for early PPD screening in perinatal populations is underexplored. METHODS: A prospective short-term longitudinal study was conducted among 64 postpartum women (32 screening-positive, 32 screening-negative). Speech recordings and socio-ecological risk factors were collected on the day before hospital discharge, and PPD screening status was determined at 42 days postpartum using the Edinburgh Postnatal Depression Scale (EPDS ≥ 10). Sixty-five acoustic features, including mel-frequency cepstral coefficients (MFCCs), zero-crossing rate, fundamental frequency, and energy parameters, were extracted and combined with seven early key risk factors. Seven machine learning classifiers were evaluated under 10-fold GroupKFold cross-validation, and model performance was assessed using AUC, sensitivity, specificity, and related metrics. The best-performing model was interpreted using SHapley Additive exPlanations (SHAP). RESULTS: Voice-only models substantially outperformed risk-only models across all classifiers. The fusion of voice and risk factors achieved the highest AUC of 0.871 (95% CI, 0.818-0.920) for the random forest model; however, this improvement over the best voice-only model was not statistically significant (p = 0.147). SHAP analysis revealed that acoustic features dominated model predictions, with MFCC and energy-related features ranking highest. Despite suboptimal probability calibration, decision curve analysis demonstrated positive net benefit across a clinically relevant range of threshold probabilities. CONCLUSION: Voice acoustic features collected before hospital discharge show potential as objective markers for early PPD screening, although further validation in larger, independent cohorts is required before clinical application.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

A multimodal feature fusion model integrating voice acoustic features and early key risk factors for early screening of postpartum depression: a prospective short-term longitudinal observational study. — 科研速览 Science Skim