Rethinking Procedural Audio Pre-training: Source Scaling and Objective Adaptation

📅 2026-09-14
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究探讨了过程音频的缩放原则及训练选择,通过控制源分离缩放类型,使用FDSL和AudioMAE实验,提出应根据源特性调整预训练策略。
📝 Abstract
Procedural audio has emerged as a viable source for transferable audio representation learning, but its design principles remain unclear.We revisit two questions: how a procedural source should be scaled, and whether training choices developed on natural audio should transfer unchanged to procedural data.Using a controlled source, we separate scale into formula-class coverage C and within-class rendering diversity I.Experiments with FDSL and AudioMAE show that these two forms of scale provide different benefits and depend on the learning formulation and downstream task. A matched AudioMAE study further shows that procedural audio favors low mask ratios (10%--25%), whereas AudioSet-28K favors 50%--75%. Shared-codebook analysis reveals lower patch diversity and stronger temporal predictability in procedural audio. These results motivate source-aware procedural pre-training, where source scaling and learning configuration are considered jointly.Code is available at https://github.com/Cross-Innovation-Lab/Formula-Bank.
Problem

Research questions and friction points this paper is trying to address.

Procedural Audio
Pre-training
Source Scaling
Objective Adaptation
Transfer Learning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Procedural Audio
Source Scaling
Objective Adaptation
Mask Ratios
Temporal Predictability
🔎 Similar Papers
No similar papers found.