科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ International Journal of Advances in Data and Information Systems2026-08-01· Computer science

Generalization Analysis of YAMNet-DNN Architectures in Deepfake Audio Classification

Hakam Dzakwan Diash, Dwi Arman Prasetya, Alfan Rizaldy Pratama, Tresna Maulana Fahrudin

原始摘要(英文原文)· Original abstract
The rapid advancement of speech synthesis and voice conversion technologies has increased the risk of audio deepfake attacks, necessitating robust and generalizable detection systems. This study proposes a deepfake audio detection framework that leverages pretrained YAMNet embeddings as a feature extractor, combined with statistical aggregation (Mean, Standard Deviation, and Mean + Standard Deviation) and Deep Neural Network classifiers to form compact utterance-level representations without temporal modeling. Experiments on a public dataset show that Mean and Mean + Standard Deviation representations achieve classification accuracy of up to 99.24%, outperforming Standard Deviation–based features, while deeper architectures provide only marginal performance gains. However, cross-dataset evaluation on a newly collected primary dataset reveals a significant generalization gap, with accuracy dropping to approximately 62% under direct testing. To address this issue, fine-tuning strategies with and without layer freezing are explored under limited target-domain data conditions, where the best-performing model achieves an accuracy of 0.93 on the primary dataset and 0.91 on the secondary dataset, with macro-averaged precision, recall, and F1-score reaching 0.93 and 0.91, respectively. Further evaluation on external datasets demonstrates that the model benefits from fine-tuning even without prior exposure to these data sources, indicating its ability to generalize across different synthesis conditions. These findings confirm that controlled fine-tuning effectively balances adaptation and generalization, while also highlighting the potential of incremental learning strategies to maintain robustness against evolving deepfake generation techniques.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Generalization Analysis of YAMNet-DNN Architectures in Deepfake Audio Classification — 科研速览 Science Skim