MASQ: Mask-Aware Spatiotemporal Quantization for Unsupervised Skeleton Action Segmentation

📅 2026-08-30
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
论文提出MASQ框架,通过空间和时间维度的创新方法解决无监督骨骼动作分割中因空间遮罩和离散量化导致的动作边界不稳定问题。
📝 Abstract
Unsupervised skeleton-based temporal action segmentation is a crucial task for understanding human behavior in long untrimmed sequences. Recent approaches often rely on discrete quantization to discover action boundaries from motion representations. However, when spatial masking is introduced for representation learning, it can introduce representation ambiguity, while discrete quantization further amplifies small fluctuations in the latent space. The interaction between these two factors often leads to unstable code switching and severe temporal jitter near action boundaries.To address these limitations, we propose a novel Mask-aware Action Spatiotemporal Quantization (MASQ) framework. Our framework decouples the conflicting tasks of spatial feature inference and temporal smoothing.In the spatial dimension, we introduce a Joint-Level Structured Dropout (JLSD) mechanism that masks the entire temporal trajectory of selected joints, to encourage the model to learn discriminative inter-joint coordination patterns. In the temporal dimension, we design a mask-aware velocity loss that enforces motion consistency only on visible joints, that prevents gradient conflicts caused by masked signals and stabilizing temporal predictions. Extensive experiments on three widely used skeleton datasets, including HuGaDB, LARa, and BABEL, demonstrate that the proposed MASQ framework significantly outperforms existing state-of-the-art unsupervised methods. In particular, our model establishes a comprehensive and substantial leading advantage in the Mean over Frames accuracy.
Problem

Research questions and friction points this paper is trying to address.

Unsupervised Skeleton-based Temporal Action Segmentation
Spatial Masking
Discrete Quantization
Representation Ambiguity
Temporal Jitter
Innovation

Methods, ideas, or system contributions that make the work stand out.

Mask-Aware Spatiotemporal Quantization
Joint-Level Structured Dropout
Unsupervised Skeleton Action Segmentation
💼 Related Jobs
No related jobs found.
X
Xinyao Qin
University of Science and Technology of China, China
L
Linxiang Peng
University of Science and Technology of China, China
Y
Youbao Ye
University of Science and Technology of China, China
Di Yang
Di Yang
School of Mathematical Sciences, University of Science and Technology of China
mathematics
Jiangtao Wang
Jiangtao Wang
Coventry University, United Kingdom
AI for HealthCrowd SensingUbiquitous ComputingDigital Health