Human-Grounded Calibration for Long-Text Image-Text Congruence in Vision-Language Models

📅 2026-09-14
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出了一种轻量级校准层Congruency Score,用于解决长文本图像-文本一致性评分问题,通过映射相似性证据到有界分数来提高视觉-语言模型的解释性和性能。
📝 Abstract
Long-text image--text congruence scoring is increasingly important for vision-language systems that must evaluate whether detailed textual descriptions match visual content. However, raw similarity scores from dual-encoder models are difficult to interpret as calibrated congruence measures, especially under the modality gap between image and text embeddings. This paper proposes Congruency Score (CS), a lightweight calibration layer that maps image--text similarity evidence into a bounded score. Using DOCCI and Urban1k, we evaluate four frozen vision-language backbones and show that observed reductions in post-projection centroid distance do not uniformly improve image--text retrieval performance. Human-grounded evaluations on DOCCI further reveal a trade-off: direct post-hoc calibration preserves high association with human judgments, whereas selected projection-based configurations can reduce threshold-relevant slope and intercept distortions at the cost of retrieval performance and association strength. These results establish long-text image--text congruence scoring as a calibrated score-estimation problem, where retrieval performance, human association, and threshold calibration must be evaluated as distinct objectives. CS provides a lightweight way to expose and operationalize this separation.
Problem

Research questions and friction points this paper is trying to address.

long-text image--text congruence
calibration
vision-language models
similarity scores
modality gap
Innovation

Methods, ideas, or system contributions that make the work stand out.

Congruency Score
calibration layer
long-text image--text congruence
vision-language models
human-grounded evaluation
🔎 Similar Papers