🤖 AI Summary
This study addresses the limitation of existing semantic communication systems in accommodating personalized receiver requirements by proposing a novel framework integrating vision-language models with latent diffusion models. The proposed approach extracts and quantizes user-preference semantic tokens to enable receiver-aware encoding, alongside a joint distortion metric that balances source fidelity with individual preferences for rate-distortion optimization under bandwidth constraints. Experimental evaluations in wireless transmission scenarios demonstrate that this framework significantly outperforms state-of-the-art baselines, such as CDDM, in both source semantic consistency and personalized reconstruction quality. Consequently, the method effectively reconciles transmission efficiency with enhanced user experience, validating its efficacy in enabling personalized digital semantic communication tailored to diverse receiver needs within resource-limited environments.
📝 Abstract
Semantic communication (SC) enables bandwidth-efficient wireless image transmission, but most existing SC schemes are user-agnostic and ignore receiver-dependent semantics. To address this issue, we propose a personalized digital semantic communication (PDSC) framework that integrates a vision-language model (VLM)-based semantic encoder with a latent diffusion model (LDM)-based semantic decoder. Specifically, the semantic encoder extracts source-aware personalized semantic tokens from both the source image and the receiver's historical interactions. These tokens are vector-quantized into discrete semantic indices and further encoded into a compact fixed-length bitstream, enabling compatibility with digital transmission. At the receiver, the semantic decoder reconstructs a personalized image conditioned on the recovered semantic tokens. Furthermore, we formulate a capacity-constrained personalized semantic rate-distortion problem and introduce a semantic distortion metric that jointly characterizes source-semantic fidelity and user-preference alignment. Experiments show that PDSC achieves superior source-semantic consistency and personalization over state-of-the-art SC baselines, including CDDM and MoS, under bandwidth-limited wireless transmission.