SGWIB:Sliced Gromov-Wasserstein Information Bottleneck for Video Highlight Detection

📅 2026-09-12
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
该研究通过引入SGWIB框架解决了视频亮点检测中的时间关系保持问题,采用结构感知正则化和上下文解缠模块来学习紧凑的瓶颈表示。
📝 Abstract
Video highlight detection aims to identify temporally important segments that capture the most informative or engaging events in a video. Reliable prediction therefore requires not only discriminative segment representations but also preservation of the temporal relationships among neighboring and distant segments. The information bottleneck principle has proven effective for learning compact and task-relevant representations, yet it has not been explored for video highlight detection, and applying conventional formulations directly would overlook inter-segment relational structure and distort highlight relevant temporal organization during compression. We therefore introduce the Sliced Gromov-Monge Gap (SGMG), a structure aware regularizer that measures the excess relational distortion induced by a prescribed source-to-bottleneck mapping relative to an optimal sliced structural correspondence. Building on SGMG, we develop SGWIB, an information-bottleneck framework for single-modal video highlight detection that learns compact bottleneck representations while preserving inter-segment temporal structure. We further introduce Home-Away-Related Contextual Pseudo-Labels and a contextual disentanglement module that reduce sports-specific contextual bias by separating highlight oriented information from contextual patterns. Experiments on MrHiSum and MoSu show that SGWIB attains the best Kendall's tau, Spearman's rho, mAP@50, and mAP@30 among the compared single-modal methods on both datasets. On MrHiSum, the visual model improves the strongest previous results by 0.031, 0.031, 0.87, and 0.75 on these four metrics, respectively. These results show that structure-aware information-bottleneck regularization combined with contextual disentanglement improves segment-level highlight prediction.
Problem

Research questions and friction points this paper is trying to address.

video highlight detection
temporal relationships
information bottleneck
Innovation

Methods, ideas, or system contributions that make the work stand out.

Sliced Gromov-Monge Gap
Information Bottleneck
Contextual Disentanglement
Video Highlight Detection
🔎 Similar Papers
2024-03-05IEEE transactions on circuits and systems for video technology (Print)Citations: 0
H
Hanjuan Huang
Key Laboratory of Agricultural Machinery Intelligent Control and Manufacturing of Fujian Province University, College of Mechanical and Electrical Engineering, Wuyi University, Wuyishan 354300, China
Y
Yung-Chieh Yeh
National Taiwan University of Science and Technology, No. 43, Sec. 4, Keelung Rd., Taipei, Taiwan
Hsing-Kuo Pao
Hsing-Kuo Pao
National Taiwan University of Science and Technology
Machine learningComputer visionInformation security