SpeechAnnotator: A Context-Aware Multi-Agent Framework and Benchmark for Multidimensional Speech Annotation

📅 2026-09-09
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决多维度语音标注成本高、依赖外部服务等问题,提出SpeechAnnotator框架,利用开源工具和多智能体协作实现高效自动标注。
📝 Abstract
Recent controllable speech generation requires training data with fine-grained annotations of speaker traits, prosody, emotion, paralinguistic cues, acoustic scenes, and context. Existing workflows often rely on manual correction, paid hosted multimodal services, or fixed processing chains, which limits large-scale data processing through annotation cost, external-service dependence, or weak cross-stage recovery. We introduce SpeechAnnotator, a locally deployable, context-aware multi-agent framework built entirely from open-source models and tools. Supporting frontend modules first obtain speaker-aware segments and final segment transcripts, while prior evidence extractors attach heterogeneous segment-level cues. Three specialist agents then collaborate through shared state: the Planning Agent converts local audio evidence, speaker history, neighboring segments, and recording-level context into field-specific contracts; the Labeling Agent performs contract-guided multimodal prediction for directly observable attributes; and the Review Agent runs a bounded review loop that checks evidence support and cross-segment consistency, triggering relabeling only for unsupported or inconsistent fields. To address the fragmentation of existing evaluation resources across isolated tasks and narrow-domain test sets, we introduce SpeechAnnotator-Bench (SA-Bench), containing 8.87 hours of human-annotated audio across nine source formats, together with SpeechAnnotator-Eval (SA-Eval), which separates Timeline-Eval for speaker-aware timeline recovery, Closed-Eval for finite-set attributes, and Open-Eval for open-ended attributes. Experiments and ablations show that SpeechAnnotator provides a locally deployable alternative to commercial audio-capable systems, while the bounded review loop improves multidimensional annotation through evidence- and context-aware field-level recovery.
Problem

Research questions and friction points this paper is trying to address.

controllable speech generation
fine-grained annotations
large-scale data processing
Innovation

Methods, ideas, or system contributions that make the work stand out.

context-aware multi-agent framework
locally deployable
bounded review loop
multidimensional speech annotation
🔎 Similar Papers
No similar papers found.
Q
Qirui Zhan
Northwestern Polytechnical University, China
S
Shuiyuan Wang
Northwestern Polytechnical University, China
J
Jingbin Hu
Northwestern Polytechnical University, China
H
Haoyu Zhang
Northwestern Polytechnical University, China
X
Xiaming Ren
Northwestern Polytechnical University, China
J
Jinrui Liang
Northwestern Polytechnical University, China
C
Chaoren Yu
Northwestern Polytechnical University, China
B
Bengu Wu
yutuzhineng, China
Y
Yunxiang Chen
Shenzhen Pimei Technology Co., Ltd., Guangdong, China
H
Houdun Liu
Shenzhen Pimei Technology Co., Ltd., Guangdong, China
S
Su Feng
Shenzhen Pimei Technology Co., Ltd., Guangdong, China
Liumeng Xue
Liumeng Xue
Hong Kong University of Science and Technology
Audio Speech and Language ProcessingSpeech Generation
Lei Xie
Lei Xie
Northwestern Polytechnical University
speech processingspeech recognitionspeech synthesismultimediaartificial intelligence