🤖 AI Summary
Mainstream cover song identification models exhibit severe robustness deficiencies on real-world YouTube videos, particularly degrading sharply under non-standard covers (e.g., instrumental renditions, tempo/pitch-modified versions).
Method: We first systematically model the diversity and perturbation patterns of covers in online video, proposing a taxonomy that characterizes network-induced cover variations; we further construct CoverYouTube—the first video-platform-oriented, multimodally uncertainty-sampled benchmark for cover robustness evaluation.
Contribution/Results: Experiments reveal that state-of-the-art models (e.g., DeepAudio, CoverHunter) suffer a 37% average drop in ranking accuracy on CoverYouTube versus the SecondHandSongs benchmark. We identify five most challenging cover types—including instrumental and karaoke versions—exhibiting the worst retrieval performance. This work exposes the overestimation of generalization capability by existing benchmarks and establishes a new standard and data foundation for rigorous model evaluation and improvement.
📝 Abstract
Recent advances in cover song identification have shown great success. However, models are usually tested on a fixed set of datasets which are relying on the online cover song database SecondHandSongs. It is unclear how well models perform on cover songs on online video platforms, which might exhibit alterations that are not expected. In this paper, we annotate a subset of songs from YouTube sampled by a multi-modal uncertainty sampling approach and evaluate state-of-the-art models. We find that existing models achieve significantly lower ranking performance on our dataset compared to a community dataset. We additionally measure the performance of different types of versions (e.g., instrumental versions) and find several types that are particularly hard to rank. Lastly, we provide a taxonomy of alterations in cover versions on the web.