Institution profile

Earth Species Project

Academic institutionnorthamerica · us
Official website
Research library4linked papers
Opportunities0open roles
Selected work

Representative Papers

Beyond task performance: Decoding bioacoustic embeddings with speech features

Jun 12, 2026

This study addresses the opacity of acoustic features encoded in pretrained audio embeddings, which hinders their adaptation to rare species or data-scarce scenarios in bioacoustics. For the first time, it systematically evaluates multiple pretrained models across six bioacoustic datasets for their ability to encode the 88-dimensional eGeMAPS acoustic feature set. Combining linear and nonlinear regression probes with normalized mutual information analysis, the work reveals a “no free lunch” phenomenon: different models exhibit distinct feature preferences—e.g., loudness is highly recoverable (R²=0.76), whereas fundamental frequency is markedly harder (R²=0.33). Leveraging feature recoverability and species relevance, the authors propose a data-driven model selection strategy and demonstrate that concatenating embeddings from multiple models yields optimal performance, offering an interpretable and composable guideline for embedding usage in bioacoustic tasks.

0 citationsRead paper

Beyond the Baseband: Adaptive Multi-Band Encoding for Full-Spectrum Bioacoustics Classification

Apr 30, 2026

Traditional bioacoustic classification systems are constrained by a 16 kHz sampling rate, utilizing only the baseband (0–8 kHz) and neglecting the high-frequency—often ultrasonic—components prevalent in animal vocalizations. This work proposes a multi-band encoding framework that decomposes full-spectrum animal calls into multiple frequency bands and fuses them into a unified representation for classification. It presents the first systematic investigation into the efficacy of multi-band decomposition and fusion strategies for full-spectrum bioacoustic classification, revealing that specific encoder architectures can produce decorrelated band-wise embeddings that substantially enhance class separability. Experimental results across three datasets demonstrate that the proposed fused representation significantly outperforms both baseband-only and time-stretched baseline methods on two of the datasets.

0 citationsRead paper

BAGEL: Benchmarking Animal Knowledge Expertise in Language Models

Apr 17, 2026

Current large language models lack a unified and comprehensive benchmark for evaluating animal-domain knowledge under closed-book conditions without external retrieval. To address this gap, this work proposes BAGEL, the first fine-grained closed-book question-answering benchmark specifically designed for animal knowledge. BAGEL integrates multi-source heterogeneous data from bioRxiv, Global Biotic Interactions, Xeno-canto, and Wikipedia, spanning dimensions such as taxonomy, morphology, habitat, and behavior. By combining manual curation with automated question generation, BAGEL enables precise analysis across knowledge categories, taxonomic groups, and data sources. This benchmark systematically reveals the strengths and limitations of language models in biodiversity-related knowledge, offering a new platform for assessing domain-specific generalization and the reliability of downstream applications.

0 citationsRead paper

What Matters for Bioacoustic Encoding

Aug 15, 2025

Bioacoustics lacks encoders capable of learning cross-species generalizable representations from limited labeled data; existing approaches are typically species-specific, architecture- or paradigm-restricted, and inadequately evaluated. Method: We propose the first general-purpose bioacoustic encoder supporting multi-species, multi-task learning (species classification, individual identification, behavior detection), grounded in a large-scale empirical study analyzing the impact of data diversity, model architecture, and training paradigms on representation learning. Our two-stage framework combines self-supervised pretraining—using a hybrid corpus of bioacoustic and general audio—and supervised fine-tuning, benchmarked across 26 diverse datasets. Contribution/Results: The encoder achieves state-of-the-art performance across tasks, with significantly improved in-distribution and out-of-distribution generalization. All model weights are publicly released to advance research on general-purpose bioacoustic representation learning.

0 citationsRead paper
Recent publications

Latest Papers

Beyond task performance: Decoding bioacoustic embeddings with speech features

Jun 12, 2026

This study addresses the opacity of acoustic features encoded in pretrained audio embeddings, which hinders their adaptation to rare species or data-scarce scenarios in bioacoustics. For the first time, it systematically evaluates multiple pretrained models across six bioacoustic datasets for their ability to encode the 88-dimensional eGeMAPS acoustic feature set. Combining linear and nonlinear regression probes with normalized mutual information analysis, the work reveals a “no free lunch” phenomenon: different models exhibit distinct feature preferences—e.g., loudness is highly recoverable (R²=0.76), whereas fundamental frequency is markedly harder (R²=0.33). Leveraging feature recoverability and species relevance, the authors propose a data-driven model selection strategy and demonstrate that concatenating embeddings from multiple models yields optimal performance, offering an interpretable and composable guideline for embedding usage in bioacoustic tasks.

0 citationsRead paper

Beyond the Baseband: Adaptive Multi-Band Encoding for Full-Spectrum Bioacoustics Classification

Apr 30, 2026

Traditional bioacoustic classification systems are constrained by a 16 kHz sampling rate, utilizing only the baseband (0–8 kHz) and neglecting the high-frequency—often ultrasonic—components prevalent in animal vocalizations. This work proposes a multi-band encoding framework that decomposes full-spectrum animal calls into multiple frequency bands and fuses them into a unified representation for classification. It presents the first systematic investigation into the efficacy of multi-band decomposition and fusion strategies for full-spectrum bioacoustic classification, revealing that specific encoder architectures can produce decorrelated band-wise embeddings that substantially enhance class separability. Experimental results across three datasets demonstrate that the proposed fused representation significantly outperforms both baseband-only and time-stretched baseline methods on two of the datasets.

0 citationsRead paper

BAGEL: Benchmarking Animal Knowledge Expertise in Language Models

Apr 17, 2026

Current large language models lack a unified and comprehensive benchmark for evaluating animal-domain knowledge under closed-book conditions without external retrieval. To address this gap, this work proposes BAGEL, the first fine-grained closed-book question-answering benchmark specifically designed for animal knowledge. BAGEL integrates multi-source heterogeneous data from bioRxiv, Global Biotic Interactions, Xeno-canto, and Wikipedia, spanning dimensions such as taxonomy, morphology, habitat, and behavior. By combining manual curation with automated question generation, BAGEL enables precise analysis across knowledge categories, taxonomic groups, and data sources. This benchmark systematically reveals the strengths and limitations of language models in biodiversity-related knowledge, offering a new platform for assessing domain-specific generalization and the reliability of downstream applications.

0 citationsRead paper

What Matters for Bioacoustic Encoding

Aug 15, 2025

Bioacoustics lacks encoders capable of learning cross-species generalizable representations from limited labeled data; existing approaches are typically species-specific, architecture- or paradigm-restricted, and inadequately evaluated. Method: We propose the first general-purpose bioacoustic encoder supporting multi-species, multi-task learning (species classification, individual identification, behavior detection), grounded in a large-scale empirical study analyzing the impact of data diversity, model architecture, and training paradigms on representation learning. Our two-stage framework combines self-supervised pretraining—using a hybrid corpus of bioacoustic and general audio—and supervised fine-tuning, benchmarked across 26 diverse datasets. Contribution/Results: The encoder achieves state-of-the-art performance across tasks, with significantly improved in-distribution and out-of-distribution generalization. All model weights are publicly released to advance research on general-purpose bioacoustic representation learning.

0 citationsRead paper