Guofu Zhang, Ziyi Li, Zhaopin Su, Ziqi Fang
The rise of personal speech data collection has enabled malicious actors to manipulate content using AI-powered tools via speech tampering techniques, such as copy-move forgery, deleting, homologous splicing, and heterologous splicing operations. These made-up things lead to fake voices, cheating, and spreading lies. While deep neural networks (DNNs) show promise for blind speech tampering detection, current methods lack cross-type generalisation due to software-specific limitations and rarely consider the classification of tampering types. To address this issue, this work presents an integrated detection framework (IDF) based on DNNs with three key components: 1) an encoder-decoder detection network generating a preliminary localisation mask (PLM) from Mel spectrogram features (MSFs); 2) an adaptive speech tampering localisation algorithm for refining the PLM; and 3) a dedicated classification network for tampering-type identification. Our IDF is trained using only the basic MSFs, ground-truth masks (GTMs), integer-class labels, binary cross-entropy loss, and sparse categorical cross-entropy loss. Supporting this work, we have produced a comprehensive dataset comprising 55,500 pairs of MSFs and GTMs for authentic and tampered speech samples derived from Chinese and English corpora through systematic tampering simulations. Experimental results demonstrate the consistent performance of our IDF framework across multiple languages and manipulation types. The model achieves average F-scores of 95.32% on Chinese, 96.02% on English, 88.81% on Spanish, and 95.75% on spoofed samples for detection tasks, while maintaining a localisation error below 0.1 seconds.