Do Spoken Language Models Hear Speech as They Read Text? Bridging Structural Gaps Between Speech and Text

📅 2026-08-24
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究针对语音和文本间结构差异导致的语义对齐问题,提出一种新框架以增强语音语言模型性能。
📝 Abstract
Spoken Language Models (SLMs) generate textual responses directly from speech, offering an alternative to cascaded systems. Despite recent advances, existing SLMs still exhibit weaker instruction-following behavior and limited generalization across diverse tasks compared to text-based language models. Our analysis shows that speech and text representations in current SLMs remain weakly aligned despite strong downstream performance, indicating that structural differences between continuous, temporally varying speech and discrete text remain insufficiently addressed. To address this, we propose a simple framework that decouples length mismatch from semantic alignment and encourages closer correspondence between speech and text representations. Experiments across multiple benchmarks demonstrate competitive performance against strong baselines, underscoring the importance of explicitly addressing structural differences between speech and text in SLM training. Our code is publicly available at https://github.com/jaykim9870/Do_SLMs_Hear_Speech_as_They_Read_Text.
Problem

Research questions and friction points this paper is trying to address.

Spoken Language Models
instruction-following behavior
generalization
speech and text representations
structural differences
Innovation

Methods, ideas, or system contributions that make the work stand out.

Spoken Language Models
structural differences
semantic alignment
decoupling length mismatch
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
H
Hyeonyu Kim
Maum AI Inc.
H
Hwayeon Kim
Maum AI Inc.
Youngwon Choi
Youngwon Choi
MAUM AI Inc.
Conversational AI
M
Myeongkyun Cho
Maum AI Inc., KAIST
H
Huu-Kim Nguyen
Atmanity Inc.