Toward More Reliable Artificial Intelligence: Reducing Hallucinations in Vision-Language Models

📅 2025-12-08
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Visual language models (VLMs) suffer from hallucination—generating factually inconsistent or image-ungrounded statements. To address this, we propose a training-free, parameter-free self-correction framework that iteratively refines model outputs via uncertainty-guided visual re-attention. Specifically, it quantifies token-level uncertainty across four dimensions—token entropy, attention dispersion, semantic consistency, and claim confidence—to dynamically identify and crop unreliable image regions, followed by response revision. The method requires no fine-tuning, external data, or architectural modifications. Evaluated on Qwen2.5-VL-7B, it reduces hallucination rate by 9.8 percentage points and improves object existence accuracy by 4.7 percentage points, outperforming existing training-free baselines. Our core contribution lies in the tight coupling of multi-dimensional uncertainty modeling with dynamic visual attention, enabling interpretable, lightweight, and reliable VLM inference without parameter updates.

Technology Category

Application Category

📝 Abstract
Vision-language models (VLMs) frequently generate hallucinated content plausible but incorrect claims about image content. We propose a training-free self-correction framework enabling VLMs to iteratively refine responses through uncertainty-guided visual re-attention. Our method combines multidimensional uncertainty quantification (token entropy, attention dispersion, semantic consistency, claim confidence) with attention-guided cropping of under-explored regions. Operating entirely with frozen, pretrained VLMs, our framework requires no gradient updates. We validate our approach on the POPE and MMHAL BENCH benchmarks using the Qwen2.5-VL-7B [23] architecture. Experimental results demonstrate that our method reduces hallucination rates by 9.8 percentage points compared to the baseline, while improving object existence accuracy by 4.7 points on adversarial splits. Furthermore, qualitative analysis confirms that uncertainty-guided re-attention successfully grounds corrections in visual evidence where standard decoding fails. We validate our approach on Qwen2.5-VL-7B [23], with plans to extend validation across diverse architectures in future versions. We release our code and methodology to facilitate future research in trustworthy multimodal systems.
Problem

Research questions and friction points this paper is trying to address.

Reduces hallucinations in vision-language models' outputs
Enables self-correction without retraining through uncertainty-guided re-attention
Improves grounding of responses in visual evidence to enhance reliability
Innovation

Methods, ideas, or system contributions that make the work stand out.

Training-free self-correction framework for VLMs
Uncertainty-guided visual re-attention with multidimensional quantification
Attention-guided cropping of under-explored image regions
💼 Related Jobs
No related jobs found.
K
Kassoum Sanogo
Department of CS, AI and Data Science, ESEO Engineering School, Angers, France
R
Renzo Ardiccioni
Associate Professor, Faculty of Law, Economy, Management, Le Mans Université, Le Mans, France