DuplexJail: Safety Alignment Breaks Under Spoken Interruption in Full-Duplex Models

📅 2026-09-08
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究探讨了全双工模型在语音中断下的安全性问题,通过固定延迟和拒绝触发两种中断方法测试模型的安全性,发现语音中断可作为攻击向量。
📝 Abstract
Full-duplex speech models accept user speech while generating responses, creating an underexplored attack surface. We introduce DuplexJail, which delivers fixed, request-independent spoken prompts through the user audio channel. We compare fixed-delay interruption after the harmful request ends with refusal-triggered interruption following a cue in the model's streaming text. Across four open-source models and 720 harmful requests from AdvBench and HarmBench, fixed-delay interruption raises whole-response attack success rates on AdvBench to 40.3% for PersonaPlex and 48.7% for PersonaPlex-RL, increases of +33.8 and +39.3 percentage points. The refusal-triggered policy reaches 35.6% and 48.6%, respectively, with all trials scored regardless of whether an interruption occurs. Selected conditions also increase FLM-Audio's harmful-response rate, while BayLing-Duplex shows decreases. These findings identify spoken interruption as a jailbreak attack vector and motivate evaluating safety throughout ongoing full-duplex interaction.
Problem

Research questions and friction points this paper is trying to address.

full-duplex speech models
spoken interruption
attack surface
safety alignment
Innovation

Methods, ideas, or system contributions that make the work stand out.

full-duplex speech models
spoken interruption
attack surface
fixed-delay interruption
refusal-triggered policy