CS-CLIP: Compositional Scene Graph-guided CLIP for Robust Compositional Reasoning

📅 2026-09-08
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究解决了视觉-语言模型在组合推理中的元素特定偏差问题,通过使用场景图引导的CLIP(CS-CLIP)方法来增强模型对组合结构的理解能力。
📝 Abstract
Vision-language models (VLMs) demonstrate strong performance across compositional reasoning benchmarks, which require reasoning over semantic perturbations of objects, attributes, relations, and their interactions. However, our controlled analysis reveals that existing compositionality-aware VLMs exhibit element-specific biases, often underperforming vanilla CLIP on certain compositional elements. To address this, we propose Compositional Scene Graph-guided CLIP (CS-CLIP), which uses scene graphs to identify compositional elements and construct structured negatives via selective masking. We further retain negatives that are most contradictory to the original caption, forcing the model to rely on compositional structure rather than surface cues. CS-CLIP achieves state-of-the-art compositional reasoning with robust performance across compositional elements. It also preserves general vision-language capabilities such as cross-modal retrieval and downstream visual reasoning, while requiring fewer training samples than prior methods.
Problem

Research questions and friction points this paper is trying to address.

Vision-language models
compositional reasoning
element-specific biases
CLIP
Innovation

Methods, ideas, or system contributions that make the work stand out.

Compositional Scene Graph
Selective Masking
Structured Negatives
Contradictory Negatives
Robust Compositional Reasoning
🔎 Similar Papers
No similar papers found.