Deep Learning-based Filtering for Video Coding: A Survey on Architectures, Algorithms, and Complexity Analysis
This work addresses the challenges of deploying deep learning-based in-loop filtering in consumer electronics—namely high computational complexity, limited memory bandwidth, and stringent power constraints—by proposing the first hardware-oriented 3D classification framework. It systematically surveys deep learning filtering approaches in video coding, covering integration strategies, exploitation of coding-side information, and lightweight network design. In alignment with recent standardization efforts in JVET’s Neural Network Video Coding (NNVC), the study provides an in-depth analysis of the trade-offs between rate-distortion performance and hardware feasibility. It delineates a clear evolutionary pathway from high-performance models toward low-power, real-time NPU-friendly architectures, identifies critical challenges such as inference latency and error propagation, and offers both theoretical insights and a practical roadmap for deploying intelligent video coding on edge devices.