Noise Adaptive Streaming Audio-Visual Speech Token Enhancement for Robust Full-Duplex Spoken Dialogue Models

📅 2026-09-08
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究提出AV-STE方法,通过结合噪音音频和唇部视频恢复语义语音标记,以提高全双工口语对话系统在噪声环境下的鲁棒性。
📝 Abstract
Full-duplex spoken dialogue systems enable simultaneous listening and speaking, but their audio-only perception often fails under background noise and overlapping speech, leading to incoherent responses. Recent audio-visual dialogue approaches show that incorporating visual cues such as lip movements improve robustness under audio corruption. However, existing approaches often adapt the large speech dialogue model itself to process visual input, requiring costly multimodal training. We propose AV-STE, a modular streaming audio-visual front-end that restores corrupted semantic speech tokens from noisy audio and lip video before they reach the speech LLM. The downstream dialogue model remains entirely frozen, preserving its pretrained conversational capabilities. When integrated with frozen Moshi, AV-STE improves average GPT-4o-judged response coherence from 1.42 to 1.91 under same-dataset speaker interference while largely preserving turn-taking behavior. Gains also transfer to out-of-domain Seamless Interaction.
Problem

Research questions and friction points this paper is trying to address.

full-duplex spoken dialogue systems
background noise
overlapping speech
audio-visual
robustness
Innovation

Methods, ideas, or system contributions that make the work stand out.

AV-STE
modular streaming audio-visual front-end
semantic speech tokens restoration
frozen dialogue model
response coherence
🔎 Similar Papers
No similar papers found.