GenPrior: Unleashing Text-to-Motion Generative Priors for Zero-Shot Skeleton-based Action Recognition

📅 2026-08-03
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the semantic-kinematic gap in zero-shot skeleton-based action recognition, which arises from the lack of geometric and physical constraints in textual prototypes. To bridge this gap, the study introduces, for the first time, prior knowledge from a pretrained text-to-action generation model and proposes a disentangled gated feature fusion mechanism to adaptively enhance textual embeddings. Kinematic-aware class prototypes are constructed by incorporating generated action sequences, and high-confidence unseen samples are mined from these prototypes to iteratively refine their representations. The proposed approach achieves state-of-the-art performance under both conventional and generalized zero-shot settings on the NTU-60, NTU-120, and PKU-MMD benchmarks, effectively aligning semantic concepts with kinematic structures.
📝 Abstract
Zero-shot skeleton-based action recognition (ZSAR) aims to recognize unseen action categories by aligning skeleton features with textual semantics. However, existing methods rely on text-derived prototypes that inherently lack geometric structure and physical constraints, resulting in a pronounced \textit{semantic-kinematic gap}. To bridge this gap, we propose \textbf{GenPrior}, the first framework to exploit generative priors from pre-trained Text-to-Motion (T2M) models for ZSAR. Specifically, we introduce Dispersion-Gated Feature Fusion, which distills kinematic prototypes and intra-class dispersion from generative motion sequences and employs a learned gating network to adaptively inject reliable structural cues into textual embeddings while suppressing synthetic artifacts. Furthermore, we propose Generative Prototype Refinement, which leverages these generation-enhanced prototypes as anchors to mine high-confidence unseen samples, calibrating class prototypes toward the true distribution and thereby unleashing strong performance gains. Extensive experiments on NTU-60, NTU-120, and PKU-MMD demonstrate that GenPrior achieves state-of-the-art performance under both zero-shot and generalized zero-shot settings. Code is available at https://github.com/jidongkuang/GenPrior.
Problem

Research questions and friction points this paper is trying to address.

zero-shot action recognition
skeleton-based action recognition
semantic-kinematic gap
text-to-motion
generative priors
Innovation

Methods, ideas, or system contributions that make the work stand out.

Text-to-Motion
Zero-Shot Action Recognition
Generative Prior
Prototype Refinement
Skeleton-based Recognition
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
J
Jidong Kuang
School of Cyber Science and Engineering, Southeast University
H
Hongsong Wang
School of Computer Science and Engineering, Southeast University; Key Laboratory of New Generation Artificial Intelligence Technology and Its Interdisciplinary Applications
Jie Gui
Jie Gui
Southeast University, China
Pattern Recognition and Machine LearningArtificial IntelligenceData MiningDeep LearningImage Processing and Computer Vis