Institution profile

Apply U

Research institution
Research library1linked papers
Opportunities0open roles
Selected work

Representative Papers

ColorBlindnessEval: Can Vision-Language Models Pass Color Blindness Tests?

Sep 23, 2025

This work addresses the robustness evaluation of vision-language models (VLMs) under color vision deficiencies. We propose ColorBlindnessEval—the first adversarial benchmark for VLMs inspired by the Ishihara color blindness test. It comprises 500 Ishihara-like images, each embedding digits 0–99 within chromatically complex patterns, and employs both binary (Yes/No) and open-ended question-answering prompts to systematically assess digit recognition across nine state-of-the-art VLMs, with human performance as reference. By adapting clinical color-vision testing paradigms to VLM evaluation, we uncover— for the first time—that VLMs suffer severe accuracy degradation and exhibit pervasive textual hallucinations under chromatic confusion. Experiments reveal that current VLMs heavily rely on superficial texture and contextual cues, lacking intrinsic color-invariant visual perception. These findings underscore the necessity of developing visually robust VLMs and highlight the value of color-robustness as a novel, clinically grounded evaluation dimension.

0 citationsRead paper
Recent publications

Latest Papers

ColorBlindnessEval: Can Vision-Language Models Pass Color Blindness Tests?

Sep 23, 2025

This work addresses the robustness evaluation of vision-language models (VLMs) under color vision deficiencies. We propose ColorBlindnessEval—the first adversarial benchmark for VLMs inspired by the Ishihara color blindness test. It comprises 500 Ishihara-like images, each embedding digits 0–99 within chromatically complex patterns, and employs both binary (Yes/No) and open-ended question-answering prompts to systematically assess digit recognition across nine state-of-the-art VLMs, with human performance as reference. By adapting clinical color-vision testing paradigms to VLM evaluation, we uncover— for the first time—that VLMs suffer severe accuracy degradation and exhibit pervasive textual hallucinations under chromatic confusion. Experiments reveal that current VLMs heavily rely on superficial texture and contextual cues, lacking intrinsic color-invariant visual perception. These findings underscore the necessity of developing visually robust VLMs and highlight the value of color-robustness as a novel, clinically grounded evaluation dimension.

0 citationsRead paper