🤖 AI Summary
This study addresses the challenge of adapting speech retrieval to diverse user intents by proposing INSPIRE, the first instruction-aware speech retrieval benchmark. INSPIRE enables dynamic definition of multidimensional retrieval criteria—including semantic, speaker, and acoustic attributes—via natural language instructions, and systematically evaluates four mainstream paradigms, including large language models and cascaded pipelines. Results reveal critical limitations under complex instructions: text-based models excel in semantic understanding but lack paralinguistic awareness, whereas speech-based models demonstrate the opposite pattern, with no single approach robustly handling all intent types. By bridging the gap in instruction-driven speech retrieval, this work establishes a foundational benchmark for advancing unified architecture research in the field.
📝 Abstract
Existing speech retrieval systems rely on fixed similarity matching and cannot adapt to diverse user intents. We introduce INSPIRE, the first benchmark for instruction-aware speech retrieval, in which natural-language instructions dynamically specify relevance criteria, including semantic content, speaker identity, speaking style, environmental sounds, and their combinations. We evaluate four retrieval paradigms: large audio-language models, cascaded pipelines, self-supervised speech models, and contrastive audio-language models. Our results reveal that no current method robustly handles all retrieval intents. Text-based approaches perform relatively better at semantic retrieval but struggle with paralinguistic attributes, while speech-based models are moderately better at capturing acoustic properties but falter at following instructions. These findings highlight the need for unified architectures capable of instruction-aware speech retrieval.