Multiscaled Multi-Head Attention-Based Video Transformer Network for Hand Gesture Recognition
Dynamic gesture recognition faces robustness bottlenecks due to inter-subject variations in pose, scale, and deformation. To address this, we propose the Multi-Scale Multi-Head Attention Video Transformer Network (MsMHA-VTN), a novel architecture featuring a pyramid-style multi-scale feature extraction module and the first multi-scale multi-head self-attention mechanism—where each attention head independently adapts to distinct spatiotemporal dimensions, enabling effective cross-scale temporal modeling. The model supports both unimodal (e.g., RGB) and multimodal (RGB-D) gesture recognition. Evaluated on NVGesture and Briareo benchmarks, MsMHA-VTN achieves state-of-the-art accuracy of 88.22% and 99.10%, respectively—substantially outperforming existing methods. These results demonstrate its effectiveness and strong generalization capability for complex dynamic sign language recognition under real-world variability.