GenzIQA: Generalized Image Quality Assessment using Prompt-Guided Latent Diffusion Models
Existing no-reference image quality assessment (IQA) methods exhibit poor cross-dataset generalization, particularly under distribution shifts such as user-generated content, synthetic imagery, and low-light conditions. To address this, we propose the first generic IQA framework leveraging the cross-attention mechanism of text-guided latent diffusion models (LDMs). Our method introduces learnable, quality-aware textual prompts and models prompt–image alignment to derive robust quality representations. Crucially, it exploits intermediate cross-attention features from the LDM denoising process—enabling zero-shot transfer to multiple benchmark datasets without fine-tuning. Extensive experiments demonstrate that our approach significantly outperforms state-of-the-art methods on diverse databases including LIVE-Youtube, KoNViD, and UHD-1. Moreover, it achieves superior out-of-distribution generalization, validating its effectiveness under substantial domain shifts. This work establishes a novel paradigm for leveraging generative model priors in blind IQA, bridging semantic understanding and perceptual quality estimation.