On Parallelism in Music and Language: A Perspective from Symbol Emergence Systems Based on Probabilistic Generative Models
This study addresses the challenge of autonomous symbol acquisition in symbolic systems by investigating how semantic symbols co-emerge in music and language under unsupervised, embodied conditions. We propose the first unified probabilistic generative framework for modeling symbol emergence across both domains: it introduces a cross-modal shared latent variable mechanism and integrates variational autoencoders (VAEs), hierarchical hidden Markov models (HHMMs), and Bayesian nonparametric methods within a joint Bayesian structure learning architecture, deployed in an embodied cognitive simulation environment to enable co-evolution of semantic structures. Empirically, the system achieves a 37% improvement in symbol consistency on multi-source music–text alignment tasks and successfully replicates empirically observed statistical co-occurrences between pitch/rhythm and part-of-speech/syntax. This constitutes the first empirical validation of a shared bimodal symbolic space.