科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Vicinagearth.2025-11-05· Computer science

ERF-BA-TFD+: a multimodal model for audio-visual deepfake detection

Leyan Wang, Jian Zhao, Xin Zhang, Xizhong Guo, Yuchen Yuan, Tianle Zhang, Jiaming Chu, Yuchu Jiang, Xu Yang, Lei Jin, Chi Zhang, Zhaofeng He

原始摘要(英文原文)· Original abstract
Abstract Deepfake detection is a critical task in identifying manipulated multimedia content. In real-world scenarios, deepfake content can manifest across multiple modalities, including audio and video. To address this challenge, we propose a novel multimodal deepfake detection model, ERF-BA-TFD+, which combines an enhanced receptive field (ERF) and audio-visual fusion. Our model processes both audio and video features simultaneously, leveraging their complementary information to improve detection accuracy and robustness. The key innovation of ERF-BA-TFD+ lies in its ability to model long-range dependencies within the audio-visual input, allowing it to better capture subtle discrepancies between real and fake content. In our experiments, we evaluate ERF-BA-TFD+ on the LAV-DF and DDL-AV datasets, which consist of segmented and full-length video clips. Unlike previous benchmarks that primarily focused on isolated segments, these datasets enable us to assess the model’s performance in a more comprehensive and realistic setting. Our method achieves state-of-the-art results on the DDL-AV dataset and outperforms most models on the LAV-DF dataset, demonstrating superior accuracy and processing speed compared to existing techniques. Additionally, the ERF-BA-TFD+ model showcased its effectiveness in the “Workshop on Deepfake Detection, Localization, and Interpretability,” Track 2: Audio-Visual Detection and Localization (DDL-AV), and won the first place in this competition.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

ERF-BA-TFD+: a multimodal model for audio-visual deepfake detection — 科研速览 Science Skim