Semantic Refinement of Universal Audio Representations through Audio-Description Alignment

📅 2026-09-08
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过音频-描述对齐方法改进预训练编码器,提升跨域分类准确性,使用线性探针和序列感知LLM读出进行评估。
📝 Abstract
Universal audio representations must preserve acoustic detail while making high-level concepts accessible across speech, music, environmental sound, and downstream models of different capacities. We study semantic refinement of an acoustically pretrained encoder by adding audio-description alignment to a foundation of BEST-RQ, reconstruction, and CTC. We compare matched control, shuffled-description, and correctly paired trajectories to distinguish correct correspondence from an extra contrastive objective. Each endpoint is frozen and evaluated with a temporal-mean linear probe and a sequence-aware LLM readout, testing whether the refined information is directly accessible and remains useful to a stronger model. Across three paired seeds, correct alignment improves domain-balanced classification by 4.66 points with the linear probe and 2.59 points with the sequence-aware LLM, with positive changes in every domain. Correct pairing accounts for 87% of the linear-probe gain, while the LLM shows its clearest correspondence-specific benefit in captioning. Dense acoustic objectives provide complementary gains under both readouts. A separate 24-layer continuation remains competitive with leading public encoders under the shared evaluator, supporting the recipe beyond the controlled study.
Problem

Research questions and friction points this paper is trying to address.

audio-description alignment
universal audio representations
semantic refinement
Innovation

Methods, ideas, or system contributions that make the work stand out.

audio-description alignment
semantic refinement
universal audio representations
BEST-RQ
L
Lejun Min
Alibaba Token Foundry; Center for Computer Research in Music and Acoustics, Stanford University
J
Junyu Dai
Alibaba Token Foundry
R
Ruichen Zheng
Alibaba Token Foundry
X
Xinyue Fan
Alibaba Token Foundry
Y
Yang Xiang
Alibaba Token Foundry
H
Huaichen Zhang
Alibaba Token Foundry
X
Xingchen Song
Alibaba Token Foundry
Yufei Shi
Yufei Shi
National University of Singapore
Vision computing
Han Zhao
Han Zhao
Alibaba
speechlarge langauge model
Xiangang Li
Xiangang Li
Unknown affiliation
speech recognitionnatural language processing