DHAuDS: A Dynamic and Heterogeneous Audio Benchmark for Test-Time Adaptation
Audio classifiers suffer from domain shift under acoustic environmental variations, yet existing test-time adaptation (TTA) studies predominantly evaluate performance under static or mismatched noise conditions, failing to model the diversity of real-world degradations. To address this, we propose DHAuDS—a novel, dynamic heterogeneous audio degradation benchmark specifically designed for audio TTA evaluation. Built upon four core datasets including UrbanSound8K-C, DHAuDS synthesizes degraded samples via dynamically modulated intensity control and multi-type noise superposition. It establishes four standardized benchmarks, introduces 14 differentiated evaluation metrics, and defines dynamic mixed-domain noise configurations. We conduct 124 reproducible experiments. As the first systematic framework for audio TTA, DHAuDS enables fair, cross-domain, and robust assessment of TTA methods under diverse, realistic audio degradations—substantially enhancing the comprehensiveness and credibility of audio model generalization evaluation.