科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ IEEE Transactions on Circuits and Systems for Video Technology2026-01-12· Computer science

Multi-Modal Cross-Attention-Guided Network for Audio-Visual Quality Evaluation via Visual Saliency and Mel-Spectrum Features

Junhao Lin, Yueli Cui, Chenli Fang, Binghong Pan, Chencheng Pan, Gangyi Jiang, Shiqing Zhang, Siwei Ma, Qi Tian

原始摘要(英文原文)· Original abstract
The quality evaluation of audio-visual (A/V) content has become increasingly critical in modern multimedia communication systems. Traditional single-modality quality evaluation methods and existing dedicated A/V quality models often fail to accurately assess the quality of A/V signals. To address this challenge, we propose a novel multi-modal cross-attention guided network specifically designed for A/V quality evaluation. By leveraging visual saliency and Mel-spectrum features, our network aims to achieve accurate and comprehensive quality evaluation. Specifically, distorted video frames are first converted into saliency maps, from which perceptually salient patches are selectively extracted and fed into a Convolutional Neural Network (CNN) for intra-frame visual feature extraction. Concurrently, the distorted audio signal is transformed into a Mel-spectrum, and time-frequency patches are extracted via sliding window techniques for CNN-based audio feature extraction. To effectively integrate these features and capture the long-term dependencies across consecutive A/V segments, we design a multi-modal cross-attention module that explicitly models complex inter-modal interactions. The resulting representations are then passed through a series of fully-connected (FC) layers for dimensionality reduction, ultimately deriving the quality score. Extensive experiments on three publicly available A/V quality datasets indicate that our metric outperforms the traditional quality metrics and newly-developed A/V quality metrics. The source code will be released at https://github.com/Jour3141/avqa.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Multi-Modal Cross-Attention-Guided Network for Audio-Visual Quality Evaluation via Visual Saliency and Mel-Spectrum Features — 科研速览 Science Skim