Improving Viewpoint-Invariance and Temporal Consistency for Action Detection

πŸ“… 2026-05-21
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This work addresses the limitations of existing action detection methods, which are often constrained by single-view training data and struggle to model fine-grained temporal dependencies. To overcome these challenges, the authors propose a two-stage action detection framework. In the first stage, virtual viewpoint augmentation is employed during training to extract viewpoint-invariant motion features. The second stage introduces a multi-scale temporal encoder based on a selective state space model, effectively integrating information across multiple viewpoints and temporal scales. This approach represents the first integration of virtual viewpoint augmentation with selective state space modeling for sequence representation. It achieves consistent and significant improvements over state-of-the-art methods across all splits of the PKU-MMD and BABEL benchmarks, while simultaneously enhancing viewpoint robustness and global temporal consistency.
πŸ“ Abstract
Viewpoint change invariance and action temporal consistency are critical aspects for the effective deployment of human action detection of untrimmed videos. Existing appearance-based video detection methods often struggle with limited viewpoint diversity during training, while motion-based detection approaches frequently fail to model fine-grained temporal relationships across consecutive motion windows. This paper introduces a novel two-stage action detection approach designed to improve both view-invariance and global temporal coherence properties. In the first stage, we extract motion features from augmented virtual viewpoints, solely used at training. Then, the second stage introduces a new view-invariant, multi-scale temporal encoder based on selective state-space sequence modelling to aggregate information across viewpoints and time scales. Experiments on PKU-MMD and BABEL benchmarks demonstrate that this approach significantly outperforms state-of-the-art methods in all considered splits. Code and trained models are available at: https://icb-vision-ai.github.io/HydraView-TAD
Problem

Research questions and friction points this paper is trying to address.

viewpoint-invariance
temporal consistency
action detection
untrimmed videos
motion modeling
Innovation

Methods, ideas, or system contributions that make the work stand out.

viewpoint-invariance
temporal consistency
state-space modeling
action detection
virtual viewpoints
πŸ”Ž Similar Papers
No similar papers found.
πŸ’Ό Related Jobs
No related jobs found.
Y
Yannick Porto
Universit Bourgogne Europe, CNRS, ICB
R
Renato Martins
Universit Bourgogne Europe, CNRS, ICB
T
Thomas Chalumeau
TEB Group, Prynel SAS
C
Cedric Demonceaux
Universit Bourgogne Europe, CNRS, ICB