🤖 AI Summary
This work addresses the challenge of analyzing contour-centric traditional music, such as Korean pansori, which relies on continuous pitch variations and lacks discrete, analyzable units. The authors propose an unsupervised method that employs a vector-quantized variational autoencoder (VQ-VAE) to automatically learn discrete tokens representing local pitch contours directly from raw audio. Robust quantization of continuous pitch movements is achieved through a novel reconstruction loss based on optimal alignment under multiple candidate time–pitch transformations. Without any labeled data, the learned tokens effectively recover expert-defined sigimsae categories and accurately distinguish between the two primary pansori modes—Gyemyeonjo and Ujo—thereby providing, for the first time, interpretable and stable corpus-based analytical units for contour-oriented musical traditions.
📝 Abstract
Computational analysis of music often relies on discrete representations, yet many musical traditions are organized around continuous pitch movement that resists segmentation into note-like units. For such traditions, the discrete units that analysis would build on are not given in advance. We address this gap by learning a vocabulary of local pitch-contour patterns directly from unlabeled audio, using a VQ-VAE that quantizes fixed-length contour segments into a finite codebook. To make the learned tokens stable across segmentation positions and small variations in timing and pitch range, we train the model with a reconstruction objective evaluated under the best alignment among a set of candidate temporal and pitch-domain transformations. Applied to Korean traditional music, the learned tokens recover information about expert-defined sigimsae categories without supervision, and in pansori individual tokens align with the two principal modes, Gyemyeonjo and Ujo, supporting their use as units for corpus-level analysis of contour-centric traditions.