What Makes a Good Layer? Assessing the Layer-Wise Intrinsic Properties of Music Foundation Models

📅 2026-08-14
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the lack of theoretical grounding for layer selection and the ineffectiveness of existing metrics in tonal tasks involving music foundation models. By systematically analyzing inter-layer properties across twelve models, we propose pitch-shift equivariance as an unsupervised proxy metric for layer selection. This measure effectively compensates for the limitations of general evaluations in capturing tonal representations. Experimental results demonstrate that the proposed metric aligns closely with downstream task performance. Furthermore, in few-shot scenarios, it matches or outperforms trainable multi-layer fusion approaches. Consequently, this work provides a reliable theoretical basis and practical guidance for applying music foundation models to tonal tasks, bridging the gap between representation analysis and effective layer utilization without reliance on labeled data.
📝 Abstract
Music foundation models are commonly used as frozen audio feature extractors, yet selecting which layer to extract from remains largely heuristic. Current practice defaults to fixed depths or multi-layer fusion, with limited understanding of why certain layers transfer better across downstream tasks or how representation quality varies with depth and pre-training paradigm. We conduct a systematic layer-wise analysis of 12 music foundation models spanning three pre-training paradigms (masked modeling, autoregressive modeling, and contrastive learning), characterizing their hidden representations through intrinsic geometric and transformation-based properties. Correlating label-free representation-quality metrics with layer-wise performance across 15 downstream tasks, we find that several metrics track layer quality for genre classification, emotion recognition, automatic tagging, and beat tracking, albeit with varying strength across tasks and pre-training paradigms. However, all metrics fail on tonal tasks such as key estimation and chord recognition, indicating that no single property serves as a general proxy for representation quality across music information retrieval tasks. To address this gap, we introduce a pitch-transposition equivariance measure that captures properties missed by these standard metrics, providing a consistent indicator of tonal quality across model families. Finally, we show that intrinsic metrics can serve as effective proxies for layer selection, matching or outperforming trainable multi-layer fusion methods, particularly in limited-data settings.
Problem

Research questions and friction points this paper is trying to address.

Music Foundation Models
Layer Selection
Representation Quality
Intrinsic Properties
Tonal Tasks
Innovation

Methods, ideas, or system contributions that make the work stand out.

Layer-wise Analysis
Pitch-transposition Equivariance
Intrinsic Metrics
Music Foundation Models
Layer Selection