科研速览 · Science Skim继续刷下去 · Keep skimming →
◇ arXiv2026-09-15· eess.AS

Language Orthogonalization for Zero-Shot Cross-Lingual Audio Deepfake Detection

Minu Kim, Ji Sub Um, Hoirin Kim

原始摘要(英文原文)· Original abstract
Audio deepfake detectors need to transfer to languages absent from training, as multilingual speech synthesis outpaces labeled anti-spoofing resources. While detectors increasingly rely on self-supervised speech models (S3Ms), these backbones encode language-dependent structure that confounds spoof cues. We address this confound through language orthogonalization, a target-free ridge map that removes S3M variation projected onto continuous language-identification (LID) embeddings. Across six languages, six S3M backbones, and all Leave-N-Out settings, it consistently reduces EER across unseen languages. Cross-lingual EER correlates with LID-space distance, where orthogonalization yields larger gains for more distant transfers.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Language Orthogonalization for Zero-Shot Cross-Lingual Audio Deepfake Detection — 科研速览 Science Skim