Visual Distortion Detection in UGC Images Using Large Multimodal Models

📅 2026-08-10
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenges of low accuracy in detecting localized visual distortions in user-generated content (UGC) images and poor generalization from synthetic to authentic scenarios (S2A). To this end, we propose VIGIL, a novel framework that, for the first time, leverages multi-layer features from a large language model (LLM) decoder as synchronized multi-level distortion detectors. We further introduce a distortion cue preservation mechanism to mitigate foreground-background ambiguity. Built upon a newly curated high-quality distortion dataset, VIGIL-140K, comprising 140,000 images, our approach achieves significant performance gains over strong baselines through multi-level feature fusion and post-processing optimization, enabling more precise localization of local distortions and substantially improved generalization to real-world scenes.
📝 Abstract
The localized depiction of perceptual quality has long been a crucial, yet underexplored, challenge in image quality assessment (IQA). Existing approaches based on large multimodal models (LMMs) predominantly rely on text-driven supervised fine-tuning (SFT). However, this training paradigm exhibits notable limitations in detection accuracy. Moreover, synthetically distorted images, which are often used as the primary training data source, show a significant generalization gap when deployed in real-world scenarios; thus, the \textbf{synthetic-to-authentic (\textit{S2A})} problem represents a critical challenge. Motivated by these issues, we propose \textbf{\textit{VIGIL}}, which leverages the LMM architecture for precise visual distortion detection. From a candidate pool of over 1000K samples, we construct the \textbf{\textit{VIGIL-140K}} training set, which consists of over 140K distorted images. These images are obtained through rigorous quality filtering and carefully crafted distortion injection, covering 8 major synthetic distortion categories. Our model leverages different layers of the large language model (LLM) decoder, treating them as \textit{multiple detectors} that perform synchronous distortion detection using multi-level features. Additionally, we retain distortion cues from predictions assigned to the non-distortion class, which helps mitigate the ambiguous foreground-background (\textit{FG-BG}) separation commonly encountered in the \textit{S2A} problem. After post-processing, our model consistently outperforms strong baselines on both in-domain synthetic distortion detection and \textit{S2A} tasks.
Problem

Research questions and friction points this paper is trying to address.

visual distortion detection
image quality assessment
synthetic-to-authentic
user-generated content
generalization gap
Innovation

Methods, ideas, or system contributions that make the work stand out.

synthetic-to-authentic (S2A)
large multimodal models
multi-level feature detection
distortion cue retention
VIGIL-140K
💼 Related Jobs
No related jobs found.
Ziheng Jia
Ziheng Jia
Shanghai Jiaotong University / Shanghai AILab
LLM and LMM on Visual Quality Assessment
Y
Yingji Liang
East China Normal University
Jiaying Qian
Jiaying Qian
Unknown affiliation
X
Xiongkuo Min
Shanghai Jiao Tong University