MP-Bench: Evaluating Voice Agents as a Multiparty Conversation Participant

📅 2026-09-11
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文通过引入MP-Bench,旨在评估语音代理在多方对话中的表现,特别是轮次理解和响应适当性,揭示了实时语音代理在此场景下的挑战。
📝 Abstract
Conversational voice agents have advanced significantly, offering increasingly natural human-machine interactions through both cascaded and end-to-end architectures. However, while recent benchmarks extensively evaluate dyadic interactions and passive audio comprehension, they largely overlook a prevalent real-world scenario: multi-party conversations. Evaluating agents in these settings is fundamentally more challenging than in dyadic interactions due to the exponentially greater conversational complexity. For voice agents to integrate seamlessly into human group dynamics, they must not only generate contextually appropriate responses but also demonstrate a nuanced understanding of open turn-taking. To address this gap, we introduce Multiparty Bench (MP-Bench), the first benchmark specifically designed to objectively evaluate conversational speech systems as active participants within multi-party contexts. MP-Bench assesses agent behavior along two primary dimensions: turn-taking awareness and response appropriateness. Additionally, we incorporate comprehension-based question-answering tasks as a complementary evaluation. By benchmarking 12 voice agents, we find that real-time voice agents stay at or below 22% on multiparty comprehension and remain near chance on multiparty turn-taking, exposing an open challenge for real-time voice agents under multiparty scenario.
Problem

Research questions and friction points this paper is trying to address.

Multiparty Conversations
Voice Agents
Turn-taking Awareness
Response Appropriateness
Benchmark
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multiparty Conversation
Turn-taking Awareness
Response Appropriateness