ProtoLIP: From Sentence-Level to Object-Level Evidence Disentanglement

📅 2026-09-14
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出ProtoLIP方法,通过组织视觉原型并利用查询依赖的路由机制,解决了视觉-语言模型中对象级证据分离不准确的问题。
📝 Abstract
Query-conditioned vision--language models enable fine-grained interpretation by revealing how visual evidence changes with textual queries. However, evidence conditioned on complete descriptions does not necessarily resolve into object-specific evidence, nor does an exposed evidence map necessarily identify the evidence that constitutes the model's prediction. Across multiple VLM architectures and independent benchmarks, we find that object-level queries often retain evidence from co-occurring objects and shared context. In this paper, we introduce \textbf{ProtoLIP}, a lightweight prototype-mediated evidence layer that organizes reusable visual prototypes into text-derived semantic families and uses query-dependent family routing to constrain which prototypes may provide evidence. Without spatial annotations or backbone retraining, ProtoLIP improves evidence localization and separation across query granularities, with localization gains transferring to independently pretrained VLMs with well-aligned patch--text representations. Despite using only text-derived weak supervision, ProtoLIP remains competitive with a spatially supervised grounding model while maintaining strong matching and competitive image--text retrieval. Crucially, ProtoLIP constructs its matching score directly from localized prototype evidence, enabling the score to be exactly decomposed into semantic-family and prototype contributions.
Problem

Research questions and friction points this paper is trying to address.

query-conditioned vision-language models
object-level evidence
evidence disentanglement
Innovation

Methods, ideas, or system contributions that make the work stand out.

ProtoLIP
prototype-mediated evidence layer
query-dependent family routing
evidence localization and separation
text-derived weak supervision
🔎 Similar Papers
No similar papers found.
Y
Yan Zhu
Department of Computer Science, Tulane University, New Orleans, LA, USA
Yongbo Chen
Yongbo Chen
Associate Professor Shanghai Jiao Tong University
SLAMplanning under uncertainty
Zhengming Ding
Zhengming Ding
Assistant Professor of Computer Science, Tulane University
Machine LearningComputer Vision
R
Rebecca Faust
Department of Computer Science, Tulane University, New Orleans, LA, USA