CRoPS: A Training-Free Hallucination Mitigation Framework for Vision-Language Models
This work addresses the susceptibility of vision-language models to hallucination during generation, which undermines their practical reliability. To mitigate this issue without requiring additional training, the authors propose a novel framework that constructs diverse hallucinatory variants by selectively removing critical textual tokens and then fuses multi-source hallucination signals through generalized contrastive decoding. This approach overcomes the limitations of existing methods that rely on overly narrow assumptions about hallucination origins. Extensive experiments demonstrate consistent performance gains across six benchmark datasets and three prominent vision-language models, achieving a 20% improvement in CHAI-R score and outperforming current state-of-the-art training-free hallucination suppression techniques.