Junxian Duan, Siyu Liu, Yiming Hao, Huaibo Huang, Ran He
Video forgery detection is challenging due to high computational demands and robustness issues in the era of generative AI. Existing methods struggle to capture temporal dynamics and subtle details, and may miss high-frequency features (e.g., textures, edges) critical for detecting forgery clues. To address this challenge, this paper proposes an innovative Dual Frequency-guided spatiotemporal feature learning (DFSL) network, which enhances forgery detection accuracy and robustness by incorporating frequency feature extraction alongside spatiotemporal feature capture. First, we introduce the Multi-Selective State-Space Module and Sequential Tri-frame Local Module to capture spatiotemporal information and temporal dynamics in videos more effectively. Then, we leverage the Dual-frequency Prompt Learning, which combines the Fourier Transform and the Wavelet Transform to extract global and local frequency information from the video. Finally, we explore a multimodal feature fusion scheme to enhance the model’s ability to capture both global dynamics and subtle traces. Experimental results demonstrate that the proposed method significantly improves forgery detection accuracy while maintaining low computational complexity and strong robustness.