From Uncertainty to Trust: Enhancing Reliability in Vision-Language Models with Uncertainty-Guided Dropout Decoding
LVLMs frequently suffer from hallucinations and unreliable outputs due to misinterpretation of visual inputs. To address this, we propose an uncertainty-guided inference-time visual token dropout method: (1) the first adaptation of dropout to the visual token level during inference; (2) decoupled modeling of epistemic and aleatoric uncertainty, with explicit focus on quantifying perceptual errors; (3) uncertainty estimation via projection of visual tokens into the text embedding space, followed by weighted masking; and (4) robust, training-free correction via multi-context masked decoding and ensemble prediction. Evaluated on CHAIR, THRONE, and MMBench, our method significantly reduces object hallucination (OH) while substantially improving output reliability and cross-scenario generation quality.