LHSDet: High-Resolution AI-Generated Image Detection via Visual Question Answering
This work addresses the limitations of existing AI-generated image detection methods, which often suffer from the loss of low-level texture details due to downsampling and exhibit limited generalization capability. To overcome these challenges, the study formulates high-resolution generated image detection as a visual question answering task and introduces a novel three-branch multimodal architecture. This architecture integrates a custom-designed low-level texture branch, a high-level semantic branch based on SigLIP2, and textual descriptions generated by BLIP-2, enabling an end-to-end differentiable detection model. A hierarchical fusion mechanism is further devised to effectively combine features across low-level, high-level, and semantic modalities. The proposed approach achieves high accuracy and strong robustness on high-resolution images synthesized by diverse diffusion and autoregressive models, significantly enhancing generalization to unseen generative models.