NVV-Locator: From Transcript Tags to Acoustic Boundaries for Fine-Grained Nonverbal Vocalization Grounding

📅 2026-09-09
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究提出NVV-Locator,通过非自回归填槽架构预测非言语发声的时间边界和类别,解决了现有方法对这些发声时间标记不足的问题。
📝 Abstract
Human speech includes nonverbal vocalizations (NVVs), such as laughter, sighs, breaths, and coughs, which convey affective and interactional information. Existing approaches typically represent NVVs as transcript-level tags, providing limited supervision for their waveform-time boundaries. We present NVV-Locator for fine-grained NVV temporal grounding. We first unify 26 NVV categories across public resources and construct large-scale timestamp-supervised training data through dual-LLM verification, transcript-guided forced alignment, and energy-based boundary refinement. We further introduce NVV-TimeBench, an expert-refined benchmark with 667 utterances and 1,094 events. NVV-Locator uses a non-autoregressive slot-filling architecture to jointly predict lexical timestamps, NVV categories, and event boundaries. On NVV-TimeBench, it achieves 71.0% Micro F1, 70.2% Macro F1, 80.4% Macro mIoU, and 59.6 ms Macro mMAE, outperforming the evaluated large audio model counterparts. Evaluation on an external corpus further demonstrates the cross-corpus generalization of NVV-Locator.
Problem

Research questions and friction points this paper is trying to address.

nonverbal vocalizations
temporal grounding
transcript-level tags
waveform-time boundaries
Innovation

Methods, ideas, or system contributions that make the work stand out.

nonverbal vocalizations
timestamp-supervised training data
non-autoregressive slot-filling architecture
NVV-TimeBench
🔎 Similar Papers
No similar papers found.
Y
Yuang Cao
ASLP@NPU, Northwestern Polytechnical University, China
Bingshen Mu
Bingshen Mu
Northwestern Polytechnical University
Speech RecognitionSpeech Understanding
Z
Zhennan Lin
ASLP@NPU, Northwestern Polytechnical University, China
G
Guojian Li
ASLP@NPU, Northwestern Polytechnical University, China
H
Haoyue Zhan
Shanghai Lingguang Zhaxian Technology, China
J
Jie Liu
Shanghai Lingguang Zhaxian Technology, China
C
Chuan Xie
Shanghai Lingguang Zhaxian Technology, China
Qiang Zhang
Qiang Zhang
University of Science and Technology of China
quantum informationquantum optics
Liumeng Xue
Liumeng Xue
Hong Kong University of Science and Technology
Audio Speech and Language ProcessingSpeech Generation
Lei Xie
Lei Xie
Northwestern Polytechnical University
speech processingspeech recognitionspeech synthesismultimediaartificial intelligence