Rethinking the Role of Spatial Mixing
This work investigates the fundamental role and necessity of spatial mixing operations in visual models. We conduct systematic ablation studies on ResNet and ConvMixer architectures, complemented by PGD-based adversarial evaluation and pixel-wise random shuffling reconstruction tests. Our findings reveal: (1) spatial mixing can be drastically simplified—retaining only random initialization or even full parameter freezing—while preserving over 98% of original ImageNet classification accuracy; (2) such simplification not only preserves performance but substantially enhances adversarial robustness (marked improvement in PGD attack accuracy) and structural recovery capability (successful reconstruction of severely shuffled pixel images). Crucially, we demonstrate for the first time that the core value of spatial mixing lies not in learning complex transformations, but in providing a lightweight, robust, and interpretable implicit structural prior—challenging prevailing assumptions about the necessity of learned spatial aggregation in vision models.