科研速览 · Science Skim继续刷下去 · Keep skimming →
◇ arXiv2026-08-23· cs.SD

AudioNoisePrints: Model-free audio watermarking using spatial correlation in flow matching TTS

Timothy Tin-Long, Jian Zhu, Aidan Pine, Mengzhe Geng

原始摘要(英文原文)· Original abstract
We present AudioNoisePrints, a training-free watermarking pipeline for flow matching and diffusion TTS models, which requires minimal extra computation during inference and does not require retraining the TTS model or reducing the generation quality. We exploited the fact that there are strong correlations between the initial Gaussian noises and the generated outputs in diffusion and flow matching models, such that a simple cosine correlation between the initial noise and the generated output can be used to perform watermaking. Moreover, we train a lightweight detector on top for more aggressive augmentations. Our method outperforms AudioSeal, a strong baseline for audio watermarking under strong augmentations. We experimented on F5TTS and other TTS and vocoder models, and concluded that they all exhibit similar spatial correlation properties, suggesting our watermarking scheme can be used for more flow-matching TTS models and even vocoders in the future.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

AudioNoisePrints: Model-free audio watermarking using spatial correlation in flow matching TTS — 科研速览 Science Skim