Improving Viewpoint-Invariance and Temporal Consistency for Action Detection
This work addresses the limitations of existing action detection methods, which are often constrained by single-view training data and struggle to model fine-grained temporal dependencies. To overcome these challenges, the authors propose a two-stage action detection framework. In the first stage, virtual viewpoint augmentation is employed during training to extract viewpoint-invariant motion features. The second stage introduces a multi-scale temporal encoder based on a selective state space model, effectively integrating information across multiple viewpoints and temporal scales. This approach represents the first integration of virtual viewpoint augmentation with selective state space modeling for sequence representation. It achieves consistent and significant improvements over state-of-the-art methods across all splits of the PKU-MMD and BABEL benchmarks, while simultaneously enhancing viewpoint robustness and global temporal consistency.