🤖 AI Summary
This work proposes a Transformer-based multi-scale dual-branch architecture to address the challenge of simultaneously modeling fine-grained multi-scale details and maintaining computational efficiency in no-reference image quality assessment. The method integrates spatial and channel-wise attention mechanisms and introduces cross-branch attention along with an adaptive pooling consistency loss to preserve spatial coherence of features across scale transformations, thereby overcoming the limitations of conventional single-scale approaches. Extensive evaluations on multiple benchmark datasets—including KonIQ-10k, LIVE, LIVE Challenge, and CSIQ—demonstrate that the proposed model significantly outperforms state-of-the-art methods, achieving notably higher correlation with human subjective quality ratings.
📝 Abstract
We present the Multi-Scale Spatial Channel Attention Network (MS-SCANet), a transformer-based architecture designed for no-reference image quality assessment (IQA). MS-SCANet features a dual-branch structure that processes images at multiple scales, effectively capturing both fine and coarse details, an improvement over traditional single-scale methods. By integrating tailored spatial and channel attention mechanisms, our model emphasizes essential features while minimizing computational complexity. A key component of MS-SCANet is its cross-branch attention mechanism, which enhances the integration of features across different scales, addressing limitations in previous approaches. We also introduce two new consistency loss functions, Cross-Branch Consistency Loss and Adaptive Pooling Consistency Loss, which maintain spatial integrity during feature scaling, outforming conventional linear and bilinear techniques. Extensive evaluations on datasets like KonIQ-10k, LIVE, LIVE Challenge, and CSIQ show that MS-SCANet consistently surpasses state-of-the-art methods, offering a robust framework with stronger correlations with subjective human scores.