🤖 AI Summary
This study systematically investigates the encoding mechanisms and recoverability of low-level acoustic attributes—namely reverberation, loudness, spectral centroid, and relative pitch—in CLAP audio embeddings. By training linear and nonlinear probing models on frozen CLAP embeddings and conducting cross-dataset and cross-model generalization analyses alongside geometric direction consistency tests, the work reveals for the first time that reverberation, loudness, and relative pitch are approximately linearly encoded, whereas the spectral centroid requires nonlinear modeling. Moreover, the linear directions associated with these attributes remain consistent across datasets and align with their corresponding textual description embeddings. These findings demonstrate that all target attributes can be reliably recovered from CLAP embeddings, confirming the model’s strong generalization capability and cross-modal consistency among eight foundational audio models.
📝 Abstract
Audio foundation models are widely adopted as general-purpose feature extractors, yet the internal structure of their learned representations remains insufficiently understood. In this work, we analyze CLAP audio embeddings through a probing framework, studying the encoding of three fundamental perceptual dimensions: reverberation (RT60), loudness (LUFS), and spectral content, measured via spectral centroid (SC) and relative pitch (RP). Probes of increasing complexity are trained to predict each attribute from frozen embeddings across five datasets spanning noise, speech, monophonic musical notes, and music mixtures. Our primary finding is that all of these attributes are reliably recoverable from the CLAP embedding space across the examined datasets. Within this global picture, two encoding regimes emerge: RT60, LUFS, and RP are approximately linearly encoded, while SC requires non-linear probes. Both regimes generalize across eight additional audio foundation models, with the notable exception that amplitude-invariant architectures discard loudness entirely by construction. The identified linear feature directions are geometrically consistent across datasets for RT60 and LUFS, while highly domain-specific for RP. Finally, we provide a qualitative demonstration of cross-modal consistency, showing that text embeddings of acoustic descriptors align geometrically with the identified RT60 feature direction.