WhAM: Towards A Translative Model of Sperm Whale Vocalization
This study addresses the need for high-fidelity, biologically consistent synthetic sperm whale click sequences (codas) to advance modeling of their social communication mechanisms. Method: We propose the first Transformer-based generative framework tailored to non-human vocalizations, built upon the music pre-trained model VampNet. Leveraging transfer learning and a novel masked phoneme modeling–autoregressive joint decoding strategy, the model enables cross-modal synthesis from arbitrary audio prompts to codas. It is fine-tuned on 10,000 field-recorded codas spanning two decades. Contribution/Results: Evaluations show that generated codas significantly outperform baselines in expert perceptual ratings and Fréchet Audio Distance. The model also demonstrates strong representational capacity in rhythm structure identification, social unit attribution, and vowel-analogy classification. This work pioneers deep generative modeling for cetacean acoustic synthesis, establishing a new paradigm for bioacoustic research and conservation.