Hierarchical Semantic-Constrained Heterogeneous Graph for Audio-Visual Event Localization

📅 2026-06-05
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenges in open-vocabulary audio-visual event localization, where unseen categories suffer from insufficient supervision, leading to weak cross-modal consistency across multiple temporal scales and semantic disconnection between clip-level and video-level representations. To tackle these issues, the authors propose a heterogeneous hierarchical graph structure that models audio-visual clips and their corresponding video-level nodes in Euclidean space, capturing intra-modal dynamics through multi-directional temporal edges. A dual-threshold gating mechanism is introduced to fuse high-confidence cross-modal information, while bidirectional semantic constraints between clips and videos enhance coherence. Furthermore, multi-level representations and textual prototypes are jointly embedded into hyperbolic space, where a hierarchical entailment regularization loss explicitly encodes semantic relationships across levels. The proposed method achieves state-of-the-art performance on the OV-AVEL benchmark, with ablation studies confirming the contribution of each component.
📝 Abstract
Open-vocabulary audio-visual event localization (OV-AVEL) jointly models audio-visual cues to recognize and temporally localize events, including categories unseen during training. Existing methods primarily learn joint audio-visual representations in Euclidean space, but still face two significant challenges. First, the lack of supervision signals for unseen categories makes it difficult to maintain audio-visual consistency across multiple temporal scales. Second, the lack of hierarchical constraints between segment- and video-level semantics prevents the model from establishing semantic consistency across different levels. To address these challenges, we propose a hierarchical semantic constrained heterogeneous graph (HSCHG) for audio-visual event localization framework. We first construct a heterogeneous hierarchical graph in Euclidean space, which includes audio and visual segment nodes and their corresponding video-level nodes. We use multi-directional temporal edges to capture complete temporal information within each modality. Simultaneously, we employ a dual-threshold filtering gated fusion strategy, introducing cross-modal information only when the alignment confidence is high. Furthermore, we introduce bidirectional semantic constraints between segment- and video-level representations to achieve semantic consistency across different levels. Based on this, we map the multi-level audio-visual representations and text prototypes uniformly into hyperbolic space. We use a hierarchical entailment regularization loss to characterize the hierarchical relationships between videos and segments. Extensive experimental results show that our method outperforms existing methods on the OV-AVEL benchmark. Ablation studies further validate the effectiveness of our method.
Problem

Research questions and friction points this paper is trying to address.

audio-visual event localization
open-vocabulary
hierarchical semantics
semantic consistency
unseen categories
Innovation

Methods, ideas, or system contributions that make the work stand out.

hierarchical semantic constraints
heterogeneous graph
hyperbolic space
open-vocabulary audio-visual event localization
cross-modal fusion
Z
Zhe Yang
Faculty of Computing, Harbin Institute of Technology, Harbin 150001, China
R
Ruyi Zhang
Faculty of Computing, Harbin Institute of Technology, Harbin 150001, China
H
Hongtao Chen
Faculty of Computing, Harbin Institute of Technology, Harbin 150001, China
Wenrui Li
Wenrui Li
Assistant Professor, University of Connecticut
StatisticsNetwork scienceBiostatistics
H
Hengyu Man
Faculty of Computing, Harbin Institute of Technology, Harbin 150001, China
Wangmeng Zuo
Wangmeng Zuo
School of Computer Science and Technology, Harbin Institute of Technology
Computer VisionImage ProcessingGenerative AIDeep LearningBiometrics
Xiaopeng Fan
Xiaopeng Fan
Professor, Harbin Institute of Technology
Video/ImageWireless