VADER: Adaptive Debiasing for Hallucination Mitigation in Video Large Language Models

πŸ“… 2026-08-09
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This work addresses the prevalent issue of hallucinated responses in video large language models (VideoLLMs) during open-ended video understanding, where generated answers often lack grounding in visual evidence. To mitigate this without requiring additional training, the authors propose an adaptive, training-free debiasing framework that dynamically reweights visual evidence to suppress hallucinations. The approach integrates cross-layer diagnosis of vision–text evidence flow, pre-softmax attention redistribution, masking of high-importance visual tokens, and contrastive decoding to enable fine-grained, frame-wise visual focus intervention. Experimental results demonstrate that the framework significantly improves event-level localization accuracy and temporal consistency across multiple VideoLLMs, achieving a 72.60% accuracy on the EventHallusion benchmark with LLaVA-Video-7B.
πŸ“ Abstract
Large vision-language models (LVLMs) have demonstrated strong performance in open-ended video understanding, yet they remain prone to fluent responses unsupported by video evidence. Existing training-free methods typically apply a globally fixed visual intervention or construct a contrastive branch through input perturbation. The former cannot accommodate video-dependent fusion paths, while the latter can be compensated by cross-frame redundancy. We therefore propose Video-Adaptive Debiasing via Evidence Reweighting (VADER), a training-free framework with two complementary modules. Visual Focus Reallocation (VFR) automatically instantiates an intervention policy for each video-question input: it diagnoses layer-wise visual-to-text evidence flow, determines where to intervene, and derives how strongly to reallocate pre-softmax attention from system-token to video-token blocks. Selective Evidence Erasure (SEE) independently masks high-importance visual tokens in every frame, constructing a prior-biased branch that is difficult to compensate through neighboring frames. Contrastive decoding then down-weights predictions that remain confident after selective evidence erasure. Across multiple VideoLLMs, VADER yields substantial improvements on event-level grounding and temporal consistency; on LLaVA-Video-7B, it reaches 72.60% accuracy on EventHallusion.
Problem

Research questions and friction points this paper is trying to address.

hallucination mitigation
video large language models
evidence grounding
temporal consistency
visual-language alignment
Innovation

Methods, ideas, or system contributions that make the work stand out.

adaptive debiasing
evidence reweighting
visual focus reallocation
selective evidence erasure
contrastive decoding
πŸ”Ž Similar Papers
D
Dong Xing
Changchun Institute of Optics, Fine Mechanics and Physics, Chinese Academy of Sciences; University of Chinese Academy of Sciences
J
Jiaxin Chen
H
Hang Yang
Changchun Institute of Optics, Fine Mechanics and Physics, Chinese Academy of Sciences; University of Chinese Academy of Sciences
P
Peixun Liu
Changchun Institute of Optics, Fine Mechanics and Physics, Chinese Academy of Sciences; University of Chinese Academy of Sciences
Qiushi Yang
Qiushi Yang
City University of Hong Kong
computer visionmulti-modal learningdeep learningmedical image analysis
Y
Yuqing Wang
Changchun Institute of Optics, Fine Mechanics and Physics, Chinese Academy of Sciences; University of Chinese Academy of Sciences