Illusion-Aware Visual Preprocessing and Anti-Illusion Prompting for Classic Illusion Understanding in Vision-Language Models
This work addresses the susceptibility of vision-language models (VLMs) to memory bias when interpreting classic optical illusions, which often leads them to overlook genuine perceptual discrepancies. To mitigate this, the authors propose a training-free anti-illusion framework that uniquely integrates type-specific image preprocessing—such as edge extraction, color isolation, morphological operations, and reference line superimposition—with qualitative-comparison-oriented prompt engineering. A multi-vote ensemble strategy is further introduced to steer the model toward attending to actual visual cues. Evaluated on the official test set of the CVPR 2026 DataCV Challenge, the method achieves an accuracy of 90.48%, with a human-verified subset reaching 98.41%, securing second place and demonstrating a significant advance in VLMs’ capacity to reason about visual illusions.