Cost-efficient Active Learning for Referring Image Segmentation and Grounding

📅 2026-08-31
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文通过主动学习和生成辅助区域-文本对来解决视觉基础中收集自然语言描述的瓶颈问题,优先处理含模糊区域的图像以提高标注效率。
📝 Abstract
Collecting natural-language referring expressions along with region annotations, such as masks or boxes, is a major bottleneck in visual grounding (VG), as annotators must write descriptions that distinguish target regions from visually similar ones. We tackle this by formulating active learning (AL) for VG under the realistic setting where only raw images are available without accompanying text. Since ground-truth text is unavailable, sample selection must estimate which images contain ambiguous regions that would require discriminative referring expressions. To address this, we generate auxiliary region-text pairs using foundation models, and introduce Referred Region Ambiguity, a new acquisition function that measures whether the model's confidence collapses onto a single region or disperses across multiple candidates. It allows our method to prioritize images with strong cross-region competition, which are more informative due to their visual ambiguity. We also design a referring-expression annotation interface that helps annotators quickly focus on writing discriminative language with a few clicks. Experiments on RIS and REC benchmarks show that our AL framework consistently outperforms several AL baselines, while a user study shows up to 1.6X faster description labeling of ours.
Problem

Research questions and friction points this paper is trying to address.

Active Learning
Visual Grounding
Referring Expressions
Region Annotations
Ambiguity
Innovation

Methods, ideas, or system contributions that make the work stand out.

Active Learning
Visual Grounding
Referred Region Ambiguity
Foundation Models
🔎 Similar Papers