Beyond task performance: Decoding bioacoustic embeddings with speech features
This study addresses the opacity of acoustic features encoded in pretrained audio embeddings, which hinders their adaptation to rare species or data-scarce scenarios in bioacoustics. For the first time, it systematically evaluates multiple pretrained models across six bioacoustic datasets for their ability to encode the 88-dimensional eGeMAPS acoustic feature set. Combining linear and nonlinear regression probes with normalized mutual information analysis, the work reveals a “no free lunch” phenomenon: different models exhibit distinct feature preferences—e.g., loudness is highly recoverable (R²=0.76), whereas fundamental frequency is markedly harder (R²=0.33). Leveraging feature recoverability and species relevance, the authors propose a data-driven model selection strategy and demonstrate that concatenating embeddings from multiple models yields optimal performance, offering an interpretable and composable guideline for embedding usage in bioacoustic tasks.