🤖 AI Summary
To address the challenge that static compression mechanisms fail to adapt to time-varying links in remote inference over dynamic bandwidth-constrained channels, this paper proposes Adaptive Rate Task-Oriented Vector Quantization (ARTOVeQ). Methodologically, ARTOVeQ introduces a nested codebook architecture coupled with a progressive learning algorithm, enabling parallel multi-resolution transmission and successive refinement of inference outputs. It integrates end-to-end joint optimization, task-oriented distortion metrics, successive refinement coding, and nested vector quantization. Experimental results demonstrate that ARTOVeQ achieves near-single-rate performance across multiple bitrates while supporting a wide range of bitrate budgets. Inference quality monotonically improves with increasing transmitted bits, and end-to-end latency is significantly reduced. Overall, ARTOVeQ achieves synergistic optimization of communication efficiency and task accuracy under dynamic channel conditions.
📝 Abstract
A broad range of technologies rely on remote inference, wherein data acquired is conveyed over a communication channel for inference in a remote server. Communication between the participating entities is often carried out over rate-limited channels, necessitating data compression for reducing latency. While deep learning facilitates joint design of the compression mapping along with encoding and inference rules, existing learned compression mechanisms are static, and struggle in adapting their resolution to changes in channel conditions and to dynamic links. To address this, we propose Adaptive Rate Task-Oriented Vector Quantization (ARTOVeQ), a learned compression mechanism that is tailored for remote inference over dynamic links. ARTOVeQ is based on designing nested codebooks along with a learning algorithm employing progressive learning. We show that ARTOVeQ extends to support low-latency inference that is gradually refined via successive refinement principles, and that it enables the simultaneous usage of multiple resolutions when conveying high-dimensional data. Numerical results demonstrate that the proposed scheme yields remote deep inference that operates with multiple rates, supports a broad range of bit budgets, and facilitates rapid inference that gradually improves with more bits exchanged, while approaching the performance of single-rate deep quantization methods.