Liberating LLM Capabilities in Full-Duplex Speech Models

๐Ÿ“… 2026-05-04
๐Ÿ›๏ธ arXiv.org
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
ๆœฌๆ–‡ๆๅ‡บListen-Write-Speakๆจกๅž‹๏ผŒ้€š่ฟ‡ๆ–‡ๆœฌไผ˜ๅ…ˆ็š„ไธ‰้€š้“่Œƒๅผ่งฃๅ†ณ่ฏญ้Ÿณๆจกๅž‹ไธญไปฃ็ ็”Ÿๆˆ็ญ‰ๆ–‡ๆœฌๅŽŸ็”Ÿ่ƒฝๅŠ›ๅ—้™็š„้—ฎ้ข˜ใ€‚
๐Ÿ“ Abstract
Speech-based large language models are typically constrained to spoken replies, which limits their user-facing outputs to what can be verbalized and suppresses text-native capabilities such as code generation, structured analysis, and multi-step reasoning in realtime interaction, for tasks that require persistent, structured, and inspectable intermediate outputs. Existing work improves spoken reasoning or full-duplex turn-taking, but still treats text as a hidden intermediate state or a subordinate modality rather than a first-class output channel. We propose Listen-Write-Speak (LWS), a text-first tri-channel paradigm in which a single autoregressive LLM continuously listens to user audio, writes visible free-form text as its primary output, and speaks a realtime oral response in parallel under a shared causal attention context. This behavior is implemented entirely through a Token Schema, requiring no architectural modifications, and learned via a two-stage data pipeline that synthesizes per-second cognitive annotations consistent with the revealed input timeline. Empirically, LWS demonstrates strong full-duplex interaction on Full-Duplex-Bench, reaches 4.72 on VoiceBench AlpacaEval, achieves 92.6% writing-speaking consistency, and consistently outperforms its internal ablations on URO-Bench. These results suggest that visible writing can serve as a first-class output channel for speech interaction without sacrificing realtime responsiveness. The code and dataset are available on the project page: https://royalzhang.com/project/lws-page/.
Problem

Research questions and friction points this paper is trying to address.

large language models
speech-based
text-native capabilities
full-duplex interaction
realtime responsiveness
Innovation

Methods, ideas, or system contributions that make the work stand out.

Listen-Write-Speak
Token Schema
full-duplex interaction
๐Ÿ’ผ Related Jobs
No related jobs found.
L
Luoyuan Zhang
CUHK(SZ); ModelBest
B
Bokai Xu
ModelBest
Junbo Cui
Junbo Cui
Tsinghua University
W
Weiyue Sun
ModelBest
Y
Yingjing Xu
ModelBest
Hanyu Liu
Hanyu Liu
Key Laboratory of Material Simulation Methods and Software of MOE, Jilin University
Computational scienceHigh pressure
Y
Yuan Yao
ModelBest; Tsinghua University