Direct or Mediated? Task-Dependent Audio Information Routing in Large Audio Language Models

📅 2026-08-27
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过分析大型音频语言模型在处理拼接音频时的行为差异,揭示了自动语音识别和音频问答任务间信息传递路径的不同,指出信息利用可能是模型泛化的限制因素。
📝 Abstract
Large Audio Language Models (LALMs) have demonstrated strong performance across a wide range of audio understanding tasks. However, they are typically evaluated on single, coherent audio segments, leaving their behavior under less familiar input configurations underexplored. We study this issue through a controlled setting in which two audio segments are concatenated into a single input. Across multiple LALMs, we observe a striking task-dependent robustness gap: automatic speech recognition (ASR) remains comparatively stable, whereas audio question answering (AQA) degrades substantially. To investigate the mechanisms underlying this disparity, we analyze how audio information is routed through LALM decoders using layer-wise attention knockout. The results reveal distinct task-dependent pathways. ASR relies primarily on direct retrieval from audio tokens by answer tokens, whereas AQA depends more strongly on a mediated route in which audio information is first integrated into prompt tokens and subsequently accessed during generation. We further probe prompt-token representations under audio concatenation and find that task-relevant audio attributes remain readily decodable, particularly in middle and later decoder layers, even when AQA performance deteriorates sharply. This dissociation indicates that the failure cannot be explained by complete loss of audio information from the decoder states and is instead consistent with a downstream bottleneck in retrieving or utilizing prompt-mediated information during answer generation. Together, our findings reveal task-dependent audio information routing in LALMs and highlight information utilization as a potential limitation on their generalization.
Problem

Research questions and friction points this paper is trying to address.

Large Audio Language Models
audio information routing
task-dependent robustness
automatic speech recognition
audio question answering
Innovation

Methods, ideas, or system contributions that make the work stand out.

task-dependent routing
audio information retrieval
direct vs. mediated paths
decoder layer analysis
🔎 Similar Papers
Y
Yizhou Zhang
Graduate School of Informatics, Kyoto University, Japan
Wangjin Zhou
Wangjin Zhou
Kyoto University
Deep Learning
X
Xin Gu
WXG, Tencent, China
Y
Yichi Wang
Graduate School of Informatics, Kyoto University, Japan
W
Wei Tan
WXG, Tencent, China
Y
Yi Zhao
WXG, Tencent, China
Z
Zhi Gong
WXG, Tencent, China
Keisuke Imoto
Keisuke Imoto
Kyoto University
Acoustic Signal ProcessingSound Event Detection
Tatsuya Kawahara
Tatsuya Kawahara
Professor, School of Informatics, Kyoto University
Speech Processingspeech recognitionNatural Language Processingdialogue