Tracing Audio Grounding and Answer Selection in Audio LLMs

📅 2026-09-03
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究探讨了音频大语言模型中音频如何真正影响答案选择的问题,通过特定训练数据和分析模型内部变化来加强音频信息的作用。
📝 Abstract
Audio Large Language Models (Audio LLMs) have advanced in audio understanding, yet they can still predict the answer by reasoning from textual cues or linguistic priors rather than the provided audio. A common remedy is to train models on data whose answers cannot be inferred from text alone. This approach can improve performance, but what changes within the model remains unclear. In this paper, we ask what must happen inside the model for the audio to actually determine the answer. Our findings are threefold. (1) Replacing the audio with silence or unrelated audio causes substantially larger performance degradation in the trained model than in the pretrained model. (2) Acoustic information most strongly shapes the model's representations of the answer choices in early-to-middle layers, while training mainly increases the influence of audio information on the final prediction in middle-to-late layers. (3) The weights learned during training have their largest impact in specific layer bands. Together, these results provide a mechanistic account of how training strengthens the use of acoustic evidence in Audio LLMs.
Problem

Research questions and friction points this paper is trying to address.

Audio LLMs
audio understanding
textual cues
linguistic priors
Innovation

Methods, ideas, or system contributions that make the work stand out.

Acoustic Information
Layer Influence
Training Impact
🔎 Similar Papers
No similar papers found.
H
Hyebin Cho
Korea Advanced Institute of Science and Technology, South Korea
S
Suho Yoo
Korea Advanced Institute of Science and Technology, South Korea
J
Jihoo Jung
Korea Advanced Institute of Science and Technology, South Korea
Joon Son Chung
Joon Son Chung
KAIST
Machine learningspeech processingcomputer vision