FLM: Frequency-Aware Language Models for Generative Image Compression

📅 2026-08-26
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决生成模型在图像压缩中导致的纹理和语义细节偏差问题,提出FLM模型,通过频率域概率建模提高压缩效率并保持确定性重建。
📝 Abstract
Generative models have significantly improved the performance ceiling of image lossy compression at low bitrates by exploiting learned priors. However, the generated textures and semantic details may deviate from the source content, thereby affecting the fidelity of image reconstruction. To solve these challenges, we propose FLM, a frequency-aware language model that improves compression efficiency through frequency-domain probabilistic modeling while retaining deterministic reconstruction. At the encoder, the input image is transformed into quantized DCT coefficients, which are organized into discrete sequences using macroblock-based coefficient tokenization. FLM then performs next-coefficient prediction to autoregressively estimate token-wise conditional probability distributions for arithmetic coding, thereby generating a compact bitstream. At the decoder, the LLM and arithmetic decoder jointly recover the frequency-domain data, followed by inverse transformations for image reconstruction. A task-specific frequency-domain dataset and a two-stage fine-tuning strategy are further developed to enable the model to operate across multiple bitrate settings. FLM is a versatile compressor that is compatible with both lossy compression and lossless JPEG recompression frameworks. Experiments show that FLM exceeds conventional and generative lossy compression methods in rate-distortion performance. FLM achieves BD-PSNR gains of 3.30 dB, 3.83 dB, and 3.80 dB than JPEG baseline on Kodak, Tecnick, and CLIC2020, respectively. Better qualitative quality of FLM can be achieved in improving semantically high fidelity and suppressing blocking artifacts. FLM is also validated to be applicable to the lossless recompression task with competitive performance.
Problem

Research questions and friction points this paper is trying to address.

generative models
image lossy compression
fidelity of image reconstruction
frequency-domain probabilistic modeling
deterministic reconstruction
Innovation

Methods, ideas, or system contributions that make the work stand out.

frequency-aware language model
quantized DCT coefficients
autoregressive prediction
two-stage fine-tuning
lossy and lossless compression
🔎 Similar Papers
No similar papers found.
J
Jiarun Chen
School of Electronic Information and Communications, Huazhong University of Science and Technology, Wuhan 430074, China
K
Kejun Wu
School of Electronic Information and Communications, Huazhong University of Science and Technology, Wuhan 430074, China
Li Li
Li Li
University of Science and Technology of China
video compressionpoint cloud compressionvolumetric data compression
C
Chengtao Cai
College of Intelligent Systems Science and Engineering, Harbin Engineering University, Harbin 150001, China
Zhengguo Li
Zhengguo Li
IEEE Fellow, Senior Principal Scientist, Institute for Infocomm Research
Video codingPhysics-guided AIComputational photographySensor fusionSwitched control
C
Chia-Wen Lin
Department of Electrical Engineering and the Institute of Communications Engineering, National Tsing Hua University, Hsinchu 30013, Taiwan