🤖 AI Summary
This work addresses the challenges of deploying deep learning-based in-loop filtering in consumer electronics—namely high computational complexity, limited memory bandwidth, and stringent power constraints—by proposing the first hardware-oriented 3D classification framework. It systematically surveys deep learning filtering approaches in video coding, covering integration strategies, exploitation of coding-side information, and lightweight network design. In alignment with recent standardization efforts in JVET’s Neural Network Video Coding (NNVC), the study provides an in-depth analysis of the trade-offs between rate-distortion performance and hardware feasibility. It delineates a clear evolutionary pathway from high-performance models toward low-power, real-time NPU-friendly architectures, identifies critical challenges such as inference latency and error propagation, and offers both theoretical insights and a practical roadmap for deploying intelligent video coding on edge devices.
📝 Abstract
As Ultra-High-Definition (UHD) displays and immersive media services become ubiquitous in the Internet of Things (IoT) and Consumer Electronics (CE) sectors, including 8K display and mobile devices, the demand for high-efficiency video coding is unprecedented. While Deep Learning-based Filtering (DLF) has emerged as a promising solution to mitigate compression artifacts inherent in standards like High Efficiency Video Coding (HEVC/H.265) and Versatile Video Coding (VVC/H.266), its deployment in CE devices is severely constrained by computational complexity, memory bandwidth, and power consumption. To bridge the gap between academic research and practical deployment, this paper presents a comprehensive, hardware-oriented survey of DLF techniques. We propose a systematic three-dimensional taxonomy classifying methods into (1) Integration Scheme within the Video Coding, (2) Coding Information Utilization, and (3) Network Design Strategy. Unlike prior reviews, this work critically analyzes the trade-offs between Rate-Distortion (RD) performance and hardware feasibility, highlighting the evolution from heavy, performance-oriented models to lightweight, hardware-friendly architectures targeting Neural Processing Units (NPUs). Furthermore, we incorporate the latest standardization activities from the Joint Video Experts Team (JVET) on Neural Network-based Video Coding (NNVC) to provide realistic guidelines. We also identify open challenges such as real-time inference latency and error propagation, providing a roadmap toward robust, low-power intelligent video coding in next-generation CE vision endpoints.