科研速览 · Science Skim继续刷下去 · Keep skimming →
◇ arXiv2026-08-17· cs.SD

INSPIRE: A Benchmark for Instruction-Aware Speech Retrieval

Chen-An Li, Hung-yi Lee

原始摘要(英文原文)· Original abstract
Existing speech retrieval systems rely on fixed similarity matching and cannot adapt to diverse user intents. We introduce INSPIRE, the first benchmark for instruction-aware speech retrieval, in which natural-language instructions dynamically specify relevance criteria, including semantic content, speaker identity, speaking style, environmental sounds, and their combinations. We evaluate four retrieval paradigms: large audio-language models, cascaded pipelines, self-supervised speech models, and contrastive audio-language models. Our results reveal that no current method robustly handles all retrieval intents. Text-based approaches perform relatively better at semantic retrieval but struggle with paralinguistic attributes, while speech-based models are moderately better at capturing acoustic properties but falter at following instructions. These findings highlight the need for unified architectures capable of instruction-aware speech retrieval.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

INSPIRE: A Benchmark for Instruction-Aware Speech Retrieval — 科研速览 Science Skim