What Are You Listening to? Temporal Music Grounding for Audio-to-Text Large Language Models

📅 2026-08-29
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过引入时间音乐定位任务及MusicGroundingBench基准测试,评估音频-文本大模型在音乐理解上的准确性,发现特定任务训练能显著提升性能。
📝 Abstract
Large audio-language models can produce fluent and musically plausible responses, yet it often remains unclear whether those responses are grounded in the audio input. We introduce temporal music grounding, a task in which a model returns one or more time spans corresponding to a queried musical note, event, or pattern. To evaluate this capability, we present MusicGroundingBench, a controlled benchmark suite built by rendering algorithmically generated piano MIDI to audio, yielding exact symbolic-to-audio alignment. The suite comprises two subsets: MGBench-3N, which evaluates note-level grounding in clips containing up to three notes, and MGBench-2B, which evaluates structured grounding and short-form music understanding in two-bar excerpts. Experiments show that temporal music grounding remains challenging for current audio-language models, whereas task-specific training yields substantial gains. We further report exploratory evidence on the relationship between grounding supervision and music understanding. These results establish MusicGroundingBench as a controlled testbed for assessing whether audio-language models ground their responses in temporally localized musical evidence.
Problem

Research questions and friction points this paper is trying to address.

temporal music grounding
audio-language models
music understanding
Innovation

Methods, ideas, or system contributions that make the work stand out.

temporal music grounding
MusicGroundingBench
audio-language models
🔎 Similar Papers
No similar papers found.