VoiceTrace: A Benchmark and Retrieval Framework for Who-Said-What Speech Retrieval

📅 2026-09-16
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决语音检索中同时考虑说话内容和说话人的问题,提出VoiceTrace框架,结合文本和参考语音进行联合建模以提高检索准确性。
📝 Abstract
Speech retrieval has become increasingly important as spoken content continues to grow across meetings, lectures, podcasts, and videos. Existing benchmarks and models have advanced semantic search over spoken content, but largely focus on \emph{what} is said while overlooking \emph{who} says it. In many real-world scenarios, however, users need to retrieve speech based jointly on semantic content and a target speaker, where the speaker may be specified naturally through a reference speech utterance rather than a predefined identity. To address this gap, we introduce \textbf{VoiceTrace-Bench}, a benchmark for hybrid speech retrieval in which each query combines text specifying \emph{what} to retrieve with reference speech specifying \emph{who} to retrieve. This setting requires models to integrate complementary semantic and speaker information directly from heterogeneous query inputs. Motivated by the joint audio-text modeling capabilities of audio-language models (ALMs), we develop \textbf{VoiceTrace}, a two-stage retrieval framework consisting of \textbf{VoiceTrace-Emb}, an embedding model that learns unified representations for efficient large-scale retrieval, and \textbf{VoiceTrace-Reranker}, a reranking model that jointly examines each query--candidate pair for fine-grained relevance estimation. Experiments show that VoiceTrace achieves state-of-the-art performance on established semantic speech retrieval benchmarks, while substantially outperforming cascade-based approaches on VoiceTrace-Bench, demonstrating its effectiveness for both conventional semantic retrieval and the new hybrid retrieval setting.
Problem

Research questions and friction points this paper is trying to address.

speech retrieval
semantic content
target speaker
reference speech
hybrid retrieval
Innovation

Methods, ideas, or system contributions that make the work stand out.

hybrid speech retrieval
VoiceTrace-Bench
audio-language models
unified representations
relevance estimation
🔎 Similar Papers
No similar papers found.
A
Aaron Yee
Zhejiang University
F
Fengjie Lu
Zhejiang University
Jiarui Hai
Jiarui Hai
Johns Hopkins University
computer auditiongenerative modelsmusic information retrieval
C
Chenang Jiang
Zhejiang University
H
Helin Wang
Johns Hopkins University
S
Siwei Tu
Independent researcher
W
Weitao You
Zhejiang University
Lingyun Sun
Lingyun Sun
Zhejiang University
Design IntelligenceHCIArtificial IntelligenceIndustrial DesignAIGC