Which Negatives Matter? Ask Your Text Encoder: Adaptive Similarity Margins for Dense-Caption Retrieval

📅 2026-08-19
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
该研究通过引入HN-CLIP方法,利用文本编码器自身的文本-文本几何结构来构建自适应相似度边界,解决了密集字幕检索中因近似重复字幕导致的优化过早饱和问题。
📝 Abstract
Dense-caption retrieval has recently been improved by introducing segmentation, edge maps, LLM-filtered captions, and cross-modal modules into contrastive fine-tuning. However, these methods largely inherit the same InfoNCE objective, whose optimization can prematurely saturate under a strong pre-trained initialization: on dense captions, the loss falls below 10^{-3} on 80% of batches within the first epoch, while its gradient reaches exact zero in fp32 in 47% of measurements. We find that this behavior is closely related to the large number of near-duplicate captions in dense-caption benchmarks, where a few highly similar negatives remain unresolved after the easy majority has already been separated. As a remedy, we introduce HN-CLIP, which uses the text encoder's own text-text geometry to construct per-negative adaptive similarity margins. Specifically, a detached caption-similarity matrix is added to the negative logits, assigning larger margins to more similar captions without mining, synthesizing, or resampling negatives. The resulting objective requires only one caption-similarity matrix and a masked logit addition during training, with no auxiliary data, additional parameters, offline preprocessing, or inference-time overhead. Extensive experiments on four dense-caption retrieval benchmarks show that HN-CLIP improves over the strongest competitors by +2.4--+4.3 R@1 while training 2.4x faster than GOAL and 5.4x faster than StructXLIP. Moreover, the proposed objective improves all six tested fine-tuning frameworks on the in-domain benchmarks and reaches the strongest full-data baseline with only 20% of the training data.
Problem

Research questions and friction points this paper is trying to address.

Dense-Caption Retrieval
InfoNCE Objective
Pre-trained Initialization
Near-Duplicate Captions
Negative Samples
Innovation

Methods, ideas, or system contributions that make the work stand out.

Adaptive Similarity Margins
Text Encoder
Dense-Caption Retrieval
HN-CLIP
Haoyue Liu
Haoyue Liu
School of Artificial Intelligence and Automation, Huazhong University of Science and Technology
Computer VisionEvent Camera
Y
Ye Chen
XJTU-POLIMI Joint School, Xi’an Jiaotong University, Xi’an 710049, China
Zhichao Wang
Zhichao Wang
School of Science and Engineering, The Chinese University of Hong Kong, Shenzhen 518172, China
X
Xiaoying Tang
School of Science and Engineering, The Chinese University of Hong Kong, Shenzhen 518172, China; Shenzhen Future Network of Intelligence Institute (FNii-Shenzhen)