MuSP-Bench: Advanced Multimodal Benchmarking of Music Understanding across Score and Performance

📅 2026-08-28
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为探究模型能否理解乐谱和演奏,本文引入MuSP-Bench基准,包含490个问题,评估多模态大语言模型在不同条件下的表现。
📝 Abstract
Musicians commonly communicate music through scores and performances. Scores encode musical intent, while performances realize it in sound. To investigate whether models can meaningfully engage with both modalities, we introduce MuSP-Bench, a human-authored benchmark of 490 questions targeting understanding across Musical Scores and Performances. The benchmark distinguishes itself by spanning score-based, performance-based, interpretive, and long-horizon reasoning across classical piano and orchestral works. We evaluate frontier multimodal large language models under multiple input conditions. Our results show that these models struggle substantially to understand scores, while facing even greater challenges when reasoning about performance audio. The benchmark is available at https://musp.vaclis.net/.
Problem

Research questions and friction points this paper is trying to address.

musical scores
performances
multimodal understanding
reasoning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multimodal Benchmarking
Music Understanding
Musical Scores and Performances
Interpretive Reasoning
💼 Related Jobs
No related jobs found.
M
Milan Liessens Dujardin
Bryel Labs, UC Berkeley
S
Song-Ze Yu
UC Berkeley
Kevin Miao
Kevin Miao
Apple (Vision Products Group)
3D DiffusionRepresentation LearningExplainabilityML-HCIAR/VR