Finetuning Strategies for Querying Sounds by Vocal Imitation

📅 2026-08-19
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过对比学习和联合对比-三元组学习两种微调策略,解决声音效果查询中的语音模仿问题。
📝 Abstract
This technical report describes our winning submission to the AES AIMLA 2025 Challenge on querying sound effects by vocal imitation. We investigate two complementary fine-tuning strategies: contrastive learning with a frozen, pretrained CED encoder, and joint contrastive-triplet learning with semi-hard negatives using a MobileNetV3 encoder. This report has been updated for posterity to include details released after the challenge.
Problem

Research questions and friction points this paper is trying to address.

vocal imitation
sound effects
contrastive learning
Innovation

Methods, ideas, or system contributions that make the work stand out.

contrastive learning
frozen pretrained CED encoder
joint contrastive-triplet learning
semi-hard negatives
MobileNetV3 encoder
🔎 Similar Papers