Canonical Color as a Lens into Concept Decodability in Vision Encoders and VLMs

📅 2026-09-08
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过使用典型颜色作为测试案例,探究视觉编码器是否能在去除颜色信息后仍能解码典型颜色,并分析了VLM后训练对颜色解码能力的影响。
📝 Abstract
Visual encoders construct a representation of the image input for Vision-Language models. How much conceptual, as opposed to immediately visible, information does this representation contain? We use canonical color as a controlled test case to ask whether vision encoders make canonical-color information linearly accessible, even when color is removed from the input image. We construct a dataset of objects with canonical colors, and probe vision encoders for both color and object identity using color and grayscale images. We find that canonical color remains decodable from grayscale images, and is tied to predicted object identity, indicating a conceptual link. Extending this analysis to full VLMs, we find that VLM post-training can have a surprisingly large effect on color decodability in the vision encoder. Overall, canonical color provides a usefully controllable lens for tracing object-level conceptual semantic information in vision encoders and VLMs.
Problem

Research questions and friction points this paper is trying to address.

Canonical Color
Vision Encoders
Concept Decodability
Grayscale Images
Object Identity
Innovation

Methods, ideas, or system contributions that make the work stand out.

Canonical Color
Vision Encoders
Color Decodability
Grayscale Images
🔎 Similar Papers