Semantic Noise Reduction via Teacher-Guided Dual-Path Audio-Visual Representation Learning
Existing methods jointly optimize contrastive alignment and masked reconstruction objectives, which often introduces semantic noise and causes optimization interference, thereby limiting cross-modal representation learning performance. This work proposes the TG-DP framework, which decouples reconstruction and alignment tasks along separate optimization paths for the first time. Each path employs a visibility pattern tailored to its specific objective, and a teacher model is introduced to guide the organization of visible tokens in the contrastive path, reducing interference and enhancing representation quality. The proposed method achieves significant improvements in zero-shot retrieval on AudioSet—R@1 increases from 35.2% to 37.4% (video→audio) and from 27.9% to 37.1% (audio→video)—and attains state-of-the-art linear probe performance on both AS20K and VGGSound benchmarks.