科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Sensors (Basel, Switzerland)2026-09-02

Structured Prototype Learning with Feature Fusion for Sparse and Asynchronous Audio-Visual Depression Recognition.

Zhonghui Jin, Pei He, Yangming Guo, Xiaodong Wang, Aiqing Fang

原始摘要(英文原文)· Original abstract
Audio-visual depression recognition in real-world scenarios is often challenged by temporal sparsity and cross-modal asynchrony, where depression-related cues may appear only in short segments and may not align precisely across modalities. Under such conditions, global pooling or dense attention tends to dilute sparse discriminative evidence with redundant context, leading to unstable utterance-level representations. To address this issue, we propose an audio-visual depression recognition framework that integrates modality feature adaptation, bidirectional cross-modal interaction, and graph-based prototype abstraction. Specifically, heterogeneous audio and visual streams are first transformed into compatible representations, after which bidirectional cross-attention models content-dependent dependencies across modalities without requiring index-wise correspondence. The fused tokens are then interpreted as graph nodes and aggregated into a compact set of semantic prototypes through graph convolution and differentiable soft clustering. In addition, audio perturbation is introduced during training as a task-oriented regularisation strategy for partial acoustic evidence loss and temporal misalignment. Experiments on the LMVD dataset demonstrate clear improvements on the primary depression-recognition task, while auxiliary evaluations on MIntRec and CMU-MOSI suggest that the structured prototype representation is beneficial for other temporally sparse audio-visual recognition tasks. These results indicate that structured prototype learning is effective for preserving sparse depression-related cues, while training-time perturbation provides a complementary regularisation effect under asynchronous multimodal conditions.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Structured Prototype Learning with Feature Fusion for Sparse and Asynchronous Audio-Visual Depression Recognition. — 科研速览 Science Skim