Beyond Convolution: A Taxonomy of Structured Operators for Learning-Based Image Processing
Standard convolutions, due to their fixed structure, linearity, and reliance on local averaging, struggle to capture complex image characteristics such as low-rank structures, adaptive basis representations, and non-uniform spatial dependencies. This work proposes a unified taxonomy encompassing five classes of structured operators—decomposition-based, adaptive weighting, basis-adaptive, integral/kernel-based, and attention-based—and systematically analyzes their differences along key dimensions including locality, linearity, and equivariance. By leveraging techniques such as singular value/tensor decomposition, content-adaptive weighting, learnable analysis bases, position-dependent nonlinear kernels, and attention mechanisms, the study comprehensively evaluates the performance of these operators across image-to-image and image-to-label tasks. The findings clarify the respective strengths and limitations of each operator class, offering both theoretical insights and practical guidance for future research.