Quantifying the Generation Modality Gap in Speech-Text Language Models

📅 2026-09-13
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过构建统一评估套件,量化纯语音与文本-语音语言模型在生成连贯内容上的差距,并发现联合建模显著提高了语义一致性。
📝 Abstract
Pure speech language models often lag behind text and speech-text language models in generating coherent content, but this gap is difficult to quantify because speech and text systems are typically evaluated with different metrics and trained on different data. We study the speech-text modality gap in a family of spoken language models, based on flow matching for continuous acoustic feature generation. We construct a unified generation-based evaluation suite that compares speech-only, text-only, and speech-text language models trained on matched data distributions and evaluated in matched generation settings. We evaluate generated continuations along multiple dimensions: semantic coherence, measured by transcribing generated speech and scoring it with a reference language model; local phonetic structure, measured by phone n-gram distributional statistics; speaker consistency and acoustic quality; and emotion-based distributional metrics. Across datasets, we find that joint speech-text modeling substantially improves semantic coherence. However, the improvement is not uniform across metrics: phone-level metrics change only modestly, speaker similarity and predicted quality are lower for speech-text continuations, while emotion-based distributional metrics improve. Compared with larger-scale speech-only models, our speech-text model closes much of the scaling gap in transcript-based semantic coherence, suggesting that text provides an efficient semantic training signal for spoken language modeling.
Problem

Research questions and friction points this paper is trying to address.

speech language models
text and speech-text language models
coherent content generation
modality gap
evaluation metrics
Innovation

Methods, ideas, or system contributions that make the work stand out.

unified generation-based evaluation suite
joint speech-text modeling
semantic coherence improvement
scaling gap reduction
🔎 Similar Papers