Boundary Voting Network for Ambiguity-Aware Timestamp-Supervised Action Segmentation

📅 2025-11-01
🏛️ IEEE transactions on circuits and systems for video technology (Print)
📈 Citations: 4
Influential: 0
📄 PDF
🤖 AI Summary
本文提出边界投票网络来解决时间戳监督动作分割中的边界定位不确定性问题,通过传播全局先验知识增强局部特征以精确定位动作边界。
📝 Abstract
Timestamp-supervised action segmentation aims to segment and classify actions in untrimmed videos with a random frame annotated per action. Precisely localizing action boundaries from timestamp annotations is crucial for this setting, as it enables generating framewise pseudo-labels and applying the well-explored fully-supervised training. However, prevailing methods struggle with intrinsic uncertainty in boundary localization due to less discriminative features in action-transiting regions. This imprecise boundary estimation significantly reduces the stability and reliability of the generated pseudo-labels in ambiguous action-transiting regions, consequently resulting in performance deterioration of the trained segmentation models. In our paper, we introduce the boundary voting network that mitigates feature ambiguity by hierarchically propagating video-level global prior knowledge into local action-transiting regions. By generating key action representations as votes throughout the video and targeting action-transiting regions, all votes collaboratively contribute to action-transiting feature enhancement and boundary localization refinement. Extensive experiments demonstrate the effectiveness of our method on GTEA, 50Salads, and Breakfast datasets.
Problem

Research questions and friction points this paper is trying to address.

timestamp-supervised action segmentation
boundary localization
feature ambiguity
pseudo-labels
Innovation

Methods, ideas, or system contributions that make the work stand out.

boundary voting network
global prior knowledge
action-transiting regions
feature enhancement
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Runzhong Zhang
Runzhong Zhang
Nanyang Technological University
Computer VisionVideo Understanding
Y
Yueqi Duan
Department of Electronic Engineering, Tsinghua University, Beijing, China
Yang Chen
Yang Chen
Nanyang Technological University
3D VisionDeep Learning
W
Weipeng Hu
School of Electrical and Electronic Engineering, Nanyang Technological University, Singapore
C
Chen Cai
School of Electrical and Electronic Engineering, Nanyang Technological University, Singapore
S
Suchen Wang
Amazon, Seattle, America
Y
Yap-Peng Tan
School of Electrical and Electronic Engineering, Nanyang Technological University, Singapore