TurnBench: A Multi-Domain Benchmark for Turn-Taking Dynamics in Spoken Dialogue

📅 2026-08-25
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决自然对话中轮流发言评估不足的问题,通过构建包含30小时手标对话的多领域基准TurnBench及标准化评估协议来改进。
📝 Abstract
Speakers in natural conversation take turns speaking and listening, deciding in real time when to take, hold, or yield the floor. However, turn-taking evaluation remains limited due to the lack of a consistent, linguistically grounded evaluation protocol and hand-annotated data covering diverse conversation types. To address this, we present TurnBench, a multi-domain benchmark that pairs a 30-hour, hand-labeled corpus of dyadic human conversation with a standardized evaluation protocol for end-of-turn and interruption detection. We set conversation type as a controllable experimental variable, covering six distinct interaction styles, and triple-annotate each conversation. Benchmarking 14 heterogeneous turn-taking systems, we find end-of-turn recall stable across types, while interruption false positives are strongly type-dependent and concentrated in backchannel-dense interaction styles. Although in smooth floor transfers human listeners begin speaking a median 151 ms before the current turn ends, no current system performs equivalently without incurring excessive false positives. We release our corpus, a 104-hour training set, and a public leaderboard with an interactive dataset viewer at https://turnbench.sesame.com
Problem

Research questions and friction points this paper is trying to address.

turn-taking dynamics
evaluation protocol
hand-annotated data
Innovation

Methods, ideas, or system contributions that make the work stand out.

turn-taking dynamics
multi-domain benchmark
hand-labeled corpus
standardized evaluation protocol
interruption detection