Guoqiang Zhang, Yu Liang, Kaiyue Tian, Jiachen Yi, Hadeel Alsolai, Menglu Liu, Xiyuan Hu
Current deepfake detection methods primarily focus on exploring inter-frame inconsistencies using convolutional networks, neglecting the investigation of long-range spatiotemporal inconsistencies. Simultaneously, these methods rely on single-feature exploitation for forgery detection, resulting in limited generalization capability and robustness. To address these issues, this paper proposes a novel network that comprehensively utilizes illumination-geometric features and facial forgery trace features to excavate deepfake artifacts across multiple scales. The architecture comprises three main components: First, the Lighting-Geometric Information Capture Module (LGCM) integrates facial landmark normal vectors and illumination coefficients to construct comprehensive spatiotemporal representations. Then, the Bi-directional Multiscale Enhancement Module (BMEM) captures attention information between different frames in the spatial domain and models inter-frame discrepancy attention in the temporal domain. Furthermore, the Spatio-temporal Attention Module (STAM) mines global semantics and adaptively derives long-range spatiotemporal representations. Experimental results demonstrate that the proposed method achieves high AUC values on the four subsets of FF++ C40, Celeb-DF, and DFDC datasets, outperforming the comparative methods. Similarly, cross-forgery method detection validates the robustness and generalization capability of the proposed approach.