科研速览 · Science Skim继续刷下去 · Keep skimming →
◇ arXiv2026-08-29· eess.AS

Beyond Speech: Dual-Domain SSL Fusion for Unified All-Type Audio Deepfake Detection

Cunhang Fan, Junqin Cao, Tian Gao, Zhipeng Xie, Jun Xue, Zhao Lv, Xin Fang

原始摘要(英文原文)· Original abstract
Unified all-type audio deepfake detection aims to determine whether an input clip is real or fake when its audio type may be speech, environmental sound, singing voice, or music. Existing speech-centric or type-dependent solutions are insufficient for this setting because the test-time audio type is unknown, while the required output is still a single binary decision. To address these issues, this paper proposes a dual-domain SSL fusion method that maps heterogeneous audio into a shared binary authenticity space. EAT-large and wav2vec 2.0 XLS-R-300M are used as complementary SSL feature sources, providing broad acoustic and event-level representations as well as waveform-level, vocal, and speech-sensitive representations. Layer-wise weighted fusion integrates multi-level artifacts from different transformer depths, while token-level fusion forms a unified feature pool without enforcing frame-level alignment between the two SSL streams. The fused tokens are summarized by multi-head attentive statistics pooling and classified with a binary MLP head. With conservative speech refinement applied on top of this unified core detector, the submitted system achieves 95.58% Macro-F1 on the AT-ADD Track 2 evaluation set and ranks second in the challenge.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Beyond Speech: Dual-Domain SSL Fusion for Unified All-Type Audio Deepfake Detection — 科研速览 Science Skim