Teaching Audio Models to Reason: A Unified Framework for Source- and Layer-wise Distillation
Large audio-language models exhibit limited performance on complex reasoning tasks, primarily due to the audio–text modality gap and the absence of structured intermediate supervision. To address this, we propose a unified knowledge distillation framework that transfers symbolic reasoning capabilities from a large text-based teacher model to an audio-based student model while preserving its acoustic understanding. Our approach introduces dual-dimensional distillation—across source modalities (text and audio teachers) and across hierarchical model layers—enabling fine-grained, layer-aligned knowledge transfer. Crucially, we incorporate structured intermediate supervision signals to bridge semantic discrepancies between acoustic representations and symbolic reasoning. Experiments demonstrate substantial improvements in multi-step reasoning performance for audio models, achieving state-of-the-art results across multiple benchmarks and effectively narrowing the semantic gap between speech representation and symbolic reasoning.