MAJEPPA: Morphing and Assessing in a Unified Piano Performance Space
This work proposes the first unified self-supervised representation framework for modeling piano performance across skill levels, from beginners to virtuosos. Addressing the challenge of integrating performance generation and evaluation across diverse proficiency tiers, the method jointly learns representations of musical scores and expressive performances within a shared embedding space by combining next-token prediction, InfoNCE loss, and supervised contrastive loss. The authors construct a large-scale piano performance dataset encompassing six skill levels and six recording conditions, leveraging a pretrained MIDI autoregressive model for conditional generation and multidimensional assessment. Evaluated on the newly introduced EVPMR benchmark, the model significantly outperforms existing approaches in tasks including performance quality assessment, competition ranking prediction, error detection, and technical skill classification, demonstrating the effectiveness and generalizability of its learned representations.