Human-inspired Global-to-Parallel Multi-scale Encoding for Lightweight Vision Models
Existing lightweight vision models struggle to balance parameter count, computational cost, and performance, while inadequately modeling human visual mechanisms. Inspired by the human visual system’s tendency to process scenes holistically before focusing on details—and to retain global context even during local attention—this work proposes a Global-to-Parallel Multi-scale Encoding (GPM) mechanism and introduces H-GPE, a lightweight network architecture. H-GPE employs a Global Insight Generator (GIG) to capture holistic semantics, while parallel branches concurrently model mid-to-large scale relationships and fine-grained textures, enabling synergistic integration of global and local features. Evaluated across image classification, object detection, and semantic segmentation tasks, H-GPE consistently outperforms state-of-the-art lightweight models with significantly fewer FLOPs and parameters, achieving a superior trade-off between accuracy and efficiency.