Zero-Shot Skeleton-Based Action Anticipation

📅 2026-08-14
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limited generalizability of existing skeleton-based action prediction methods to unseen categories by formally defining the task of zero-shot skeleton action prediction. We propose a cross-modal alignment baseline model grounded in mutual information maximization, which achieves accurate early recognition of unseen actions by maximizing the mutual information between partially observed features and semantic embeddings. Furthermore, we establish a comprehensive evaluation protocol for this novel task. Extensive experiments on the NTU RGB+D dataset demonstrate that the proposed model attains superior zero-shot prediction accuracy, validating its effectiveness as a robust baseline. Collectively, this work introduces a new paradigm for action prediction research in open-world scenarios, bridging the gap between observed dynamics and semantic knowledge.
📝 Abstract
Action anticipation (AA) aims to recognize ongoing human or humanoids actions from partial observations, enabling robots to predict intentions before the actions are completed. Although skeleton-based AA offers efficiency advantages, existing approaches assume that all action classes are seen during training, which limits their deployment in real-world scenarios where novel actions inevitably arise. To address this gap, we study the new task of Zero-Shot Skeleton-Based Action Anticipation (ZS-SkAA). This task requires recognizing unseen action classes using only limited early-stage skeleton sequences, combining the challenges of partial observations, temporal dynamics, and zero-shot generalization. To establish foundational research for ZS-SkAA, we introduce:(1) A baseline model comprising a spatio-temporal feature extractor and a mutual information estimation and maximization module. This baseline model explicitly aligns partial visual features with semantic class embeddings across modalities by estimating and maximizing their mutual information, enhancing generalization to unseen classes.(2) A benchmark protocol using the NTU RGB+D dataset, which is adapted for rigorous ZS-SkAA evaluation. Experiments demonstrate the effectiveness of our model as a strong baseline for ZS-SkAA, achieving high zero-shot accuracy on NTU RGB+D. This work establishes ZS-SkAA as a vital research direction for real-world systems requiring generalization to novel actions.
Problem

Research questions and friction points this paper is trying to address.

Zero-Shot Action Anticipation
Skeleton-Based Recognition
Unseen Action Classes
Partial Observations
Innovation

Methods, ideas, or system contributions that make the work stand out.

Zero-Shot Skeleton-Based Action Anticipation
Mutual Information Maximization
Cross-Modal Alignment
Partial Observations
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
H
Hongsong Wang
School of Computer Science and Engineering, Southeast University, Nanjing 210096, China
P
Pengbo Yan
Southeast University - Monash University Joint Graduate School, SuZhou 215123, China
Y
Yang Zhang
School of Computer Science and Software Engineering, Shenzhen University, Shenzhen 518060, China
Qiuxia Lai
Qiuxia Lai
Associate Professor, Communication University of China
Computer Vision