Learning Audio-Visual Embeddings with Inferred Latent Interaction Graphs
This work addresses the challenges of spurious negative samples and missing cross-modal semantic associations in audio-visual embedding learning caused by sparse annotations. To mitigate these issues, we propose a novel learning framework that leverages soft-label prediction and an implicit interaction graph. Our approach employs a teacher–student architecture to generate reliable soft supervision signals and utilizes the GRaSP algorithm to construct a directed inter-class dependency graph. By incorporating graph-guided regularization and semantic alignment losses, the model effectively captures latent semantic dependencies among unannotated co-occurring events. Experiments on the AVE and VEGAS benchmarks demonstrate that the proposed method significantly improves mean average precision (mAP), enhancing both semantic consistency and robustness in cross-modal embeddings.