Heard but Not Heeded: Paralinguistic Information Encoding and Loss in Audio-Language Models

📅 2026-09-01
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究分析了四个开源音频语言模型对副语言信息的编码和丢失情况,使用多种方法追踪风格信息从音频编码到最终输出的过程,揭示了当前模型在利用副语言信息方面的局限。
📝 Abstract
Audio language models are designed to understand speech, yet it remains unclear whether they capture how something is said beyond what is said. We present a mechanistic analysis of paralinguistic information in four open source models, Whisper-large-v2, Qwen2-Audio-7B Instruct, Qwen2.5-Omni-7B, and Chroma-4B, using the Expresso dataset with controlled speaking styles. We combine centered kernel alignment, linear probing with leave one speaker out evaluation, open ended tone prediction, and a content prosody leakage metric to trace how style information moves from the audio encoder to the final output. All models strongly encode speaking style in the late encoder, that is, the top third of the audio encoder's layers, but this information is consistently degraded before reaching the output. The projector reshapes representation geometry without removing information, while decoders differ in how much style they preserve depending on architecture and training objective. At the output level, models fall into two behaviors. Some are content driven, where predictions depend mainly on text. Others are acoustic driven, where predictions vary with speaking style. The leakage metric quantifies this difference, and qualitative results confirm it. Overall, we identify a gap between what models encode and what they use, highlighting a key limitation in current audio language models.
Problem

Research questions and friction points this paper is trying to address.

paralinguistic information
audio language models
speaking style
encoder
output
Innovation

Methods, ideas, or system contributions that make the work stand out.

paralinguistic information
audio encoder
representation geometry
content prosody leakage
model architecture
🔎 Similar Papers
No similar papers found.
B
Bhuvan Koduru
Language Technologies Institute, Carnegie Mellon University
D
Dareen Safar B Alharthi
Language Technologies Institute, Carnegie Mellon University
R
Rita Singh
Language Technologies Institute, Carnegie Mellon University
Bhiksha Raj
Bhiksha Raj
Carnegie Mellon University
Deep LearningArtificial IntelligenceSpeech and Audio ProcessingSignal ProcessingMachine Learning