Learning Pyramid-structured Long-range Dependencies for 3D Human Pose Estimation
Modeling long-range dependencies in 3D human pose estimation remains challenging due to noise susceptibility and high model complexity. To address this, we propose the Pyramid Graph Attention (PGA) module and a lightweight Multi-Scale Graph Transformer (PGFormer). Our core contribution is the first formulation of human anatomical substructures—joints, limbs, and torso—as a pyramid-shaped, cross-scale graph, coupled with a pooling-augmented self-attention mechanism that preserves structural priors while enabling multi-granularity feature interaction. By integrating graph convolutional operations with multi-scale feature fusion, our method effectively suppresses redundancy in deep networks. Evaluated on Human3.6M and MPI-INF-3DHP, it achieves state-of-the-art accuracy (MPJPE: 41.2 mm and 89.7 mm, respectively) with a 23% reduction in parameter count, demonstrating both the efficacy and efficiency of cross-scale graph modeling for 3D pose estimation.