Trust the Model: Compact VLMs as In-Context Judges for Image-Text Data Quality

📅 2025-07-27
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
To address performance degradation in vision-language models (VLMs) caused by image-text misalignment and noisy training data, this paper proposes a lightweight, self-consistent data filtering framework. The method leverages only a fine-tuned small-scale VLM—without auxiliary modules or large language/model guidance—to jointly assess image quality, text fluency, and cross-modal semantic alignment. It employs a context-aware, end-to-end judgment mechanism, enabling efficient data purification with minimal computational overhead. Experiments demonstrate that datasets filtered by our framework substantially outperform the original noisy datasets across multiple downstream vision-language tasks—and even rival human-annotated high-quality benchmarks. These results validate the effectiveness and practicality of the novel paradigm “leveraging small VLMs to drive high-quality dataset construction.”

Technology Category

Application Category

📝 Abstract
Vision-language models (VLMs) extend the conventional large language models by integrating visual data, enabling richer multimodal reasoning and significantly broadens the practical applications of AI. However, including visual inputs also brings new challenges in maintaining data quality. Empirical evidence consistently shows that carefully curated and representative training examples often yield superior results compared to simply increasing the quantity of data. Inspired by this observation, we introduce a streamlined data filtration framework that employs a compact VLM, fine-tuned on a high-quality image-caption annotated dataset. This model effectively evaluates and filters potential training samples based on caption and image quality and alignment. Unlike previous approaches, which typically add auxiliary filtration modules on top of existing full-scale VLMs, our method exclusively utilizes the inherent evaluative capability of a purpose-built small VLM. This strategy eliminates the need for extra modules and reduces training overhead. Our lightweight model efficiently filters out inaccurate, noisy web data, improving image-text alignment and caption linguistic fluency. Experimental results show that datasets underwent high-precision filtration using our compact VLM perform on par with, or even surpass, larger and noisier datasets gathered through high-volume web crawling. Thus, our method provides a lightweight yet robust solution for building high-quality vision-language training corpora. \ extbf{Availability and implementation:} Our compact VLM filtration model, training data, utility scripts, and Supplementary data (Appendices) are freely available at https://github.com/daulettoibazar/Compact_VLM_Filter.
Problem

Research questions and friction points this paper is trying to address.

Maintaining data quality in vision-language models with visual inputs
Filtering noisy web data for better image-text alignment
Reducing training overhead by using compact VLMs for filtration
Innovation

Methods, ideas, or system contributions that make the work stand out.

Compact VLM for image-text data filtration
Fine-tuned on high-quality annotated dataset
Lightweight model eliminates extra modules
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
D
Daulet Toibazar
Humain, Riyadh, KSA
K
Kesen Wang
S
Sherif Mohamed
A
Abdulaziz Al-Badawi
A
Abdulrahman Alfulayt
P
Pedro J. Moreno