Semantic-Aware Generative Image Transmission for Resource-Constrained Visual IoT Systems

📅 2026-06-24
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenge of balancing semantic fidelity and transmission efficiency in resource-constrained visual Internet-of-Things systems by proposing a semantic-aware generative image transmission framework. The approach integrates instance segmentation–driven semantic scoring with prediction entropy–guided recoverability assessment to intelligently sample discrete VQ tokens, and leverages MaskGIT for reconstructing missing content at the edge or cloud. Spatially dispersed scheduling via Halton sequences is introduced to enhance generation quality. At a bitrate of 0.074 bpp—only 44.6% of that required by DeepJSCC/WITT—the method achieves a PSNR of 29.9 dB, while downstream detection tasks demonstrate that its semantic masking strategy significantly outperforms random masking.
📝 Abstract
Resource-constrained visual Internet of Things (IoT) systems, such as edge cameras, unmanned sensing platforms, industrial inspection nodes, and remote monitoring sensors, often need to transmit task-relevant visual evidence over low-rate wireless links to an edge/cloud service. Existing image communication methods usually compress or transmit complete global representations, leaving limited room to exploit receiver-side generative restoration. This paper proposes a semantic-aware generative image transmission framework for edge-assisted visual IoT. The image captured by an IoT visual sensor is encoded into a discrete token grid by a VQ encoder. At the IoT transmitter or nearby gateway, token recoverability, estimated from prediction entropy and local structure complexity, is fused with semantic importance obtained from instance segmentation and category-aware scoring. A spatial dispersal sampler then selects the tokens to be transmitted under a bitrate budget. The transmitter sends only the quantization indices of kept tokens and a binary mask map, while the edge/cloud receiver recovers masked tokens through MaskGIT with Halton sequence scheduling. Experiments on Kodak and VisDrone scenes under AWGN and Rayleigh channels show that the proposed method provides a flexible bitrate-quality tradeoff for narrowband visual IoT links. At 0.074 bpp, it uses 44.6% of the transmitted bits of the 0.167-bpp DeepJSCC/WITT reference while achieving 29.9 dB PSNR. A pseudo-GT downstream detection study on Kodak further shows that semantic-aware masking preserves task-relevant objects better than random masking at both 30% and 50% mask ratios.
Problem

Research questions and friction points this paper is trying to address.

Visual IoT
low-bitrate transmission
semantic-aware
generative image recovery
resource-constrained
Innovation

Methods, ideas, or system contributions that make the work stand out.

semantic-aware transmission
generative image reconstruction
token-based compression
MaskGIT
visual IoT
🔎 Similar Papers
No similar papers found.
C
Chenyang Zhang
School of Computer and Information Engineering, Tianjin Normal University, Tianjin, China
C
Changwang Liu
School of Computer and Information Engineering, Tianjin Normal University, Tianjin, China
J
Jinqi Zhu
School of Computer and Information Engineering, Tianjin Normal University, Tianjin, China
J
Jiayi Chang
School of Computer and Information Engineering, Tianjin Normal University, Tianjin, China
Y
Yuxuan Wang
School of Computer and Information Engineering, Tianjin Normal University, Tianjin, China
S
Shuqing He
School of Information Science and Engineering, Linyi University, Linyi 276000, China
J
Jia Guo
School of Computer and Information Engineering, Tianjin Normal University, Tianjin, China