What Did the MLLM Hear? Token-Level Spectro-Temporal Grounding for Audio MLLM Explainability

📅 2026-09-11
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为了解决音频MLLM生成的每个标记对应输入音频部分不明确的问题,本文提出了STAG框架,通过时间支持估计和频段相关性测量来生成频谱-时间相关图。
📝 Abstract
Audio-based Multimodal Large Language Models (MLLMs) can generate detailed natural-language descriptions of complex acoustic scenes, yet it remains unclear which parts of the input audio support each generated token. This is particularly challenging because acoustic evidence is distributed across time and frequency, and concurrent sound events may overlap temporally while occupying different spectral regions. We introduce STAG, to our knowledge the first post-hoc framework for token-level spectro-temporal grounding of captions generated by audio-based MLLMs. STAG estimates the temporal support for each generated token using target-token-specific vocabulary projections of the encoded audio representations, measures frequency-band relevance through controlled spectral occlusion, and combines the two signals into a spectro-temporal relevance map. We evaluate STAG against ten post-hoc explanation methods across four grounding benchmarks, where it achieves the best event-localization performance on every dataset, and apply it to eight audio-language backbones without parameter updates. Counterfactual deletion further shows that removing the identified evidence selectively reduces confidence in the corresponding event and frequently removes it from the regenerated caption. These results provide behavioral support for the faithfulness and selectivity of the explanations.
Problem

Research questions and friction points this paper is trying to address.

audio-based MLLMs
token-level grounding
spectro-temporal relevance
Innovation

Methods, ideas, or system contributions that make the work stand out.

Spectro-temporal grounding
Token-level explanation
Multimodal Large Language Models
STAG
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
L
Lucia Cascone
Department of Computer Science, University of Salerno, Italy
V
Valeria Fraenza
Department of Computer Science, University of Salerno, Italy
M
Michele Nappi
Department of Computer Science, University of Salerno, Italy
F
Fabio Narducci
Department of Computer Science, University of Salerno, Italy
B
Benedetto Simone
Department of Computer Science, University of Salerno, Italy