Beyond Visual Similarity: Entity-Aligned Retrieval for Knowledge-Based Visual Question Answering

📅 2026-08-19
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决视觉相似性与实体语义一致性之间的矛盾,提出KBMR方法,利用多模态大语言模型生成语义嵌入和一致性权重,提升知识库视觉问答的检索精度。
📝 Abstract
Knowledge-Based Visual Question Answering (KB-VQA) relies on retrieving external information to answer queries involving long-tail entities. However, existing retrieval pipelines predominantly employ CLIP-style dual encoders, which prioritize surface-level visual similarity over entity-level semantic alignment. This paradigm often fails when semantically identical concepts exhibit large visual variations or when distinct entities appear visually similar. To address this, we propose KBMR, the first MLLM-based embedding retriever tailored for KB-VQA. Leveraging the robust autoregressive capabilities of MLLMs, KBMR maps images into a semantic space that better preserves concept identity. To tackle the challenge of noisy supervision in Wikipedia-scale retrieval, we introduce an MLLM-based semantic discriminator that generates continuous entity-consistency weights. These weights guide a novel continuous semantic distillation objective, enabling effective hard negative sampling and soft supervision beyond rigid binary labels. Extensive experiments demonstrate that KBMR significantly outperforms CLIP baselines, yielding up to a 14.7% improvement in retrieval Recall@1 and a 9.4% gain in end-to-end VQA accuracy. Code is available at https://github.com/realHarryX/KBMR.
Problem

Research questions and friction points this paper is trying to address.

Knowledge-Based Visual Question Answering
entity-level semantic alignment
visual similarity
Innovation

Methods, ideas, or system contributions that make the work stand out.

MLLM-based embedding retriever
semantic space
entity-consistency weights
continuous semantic distillation
🔎 Similar Papers
No similar papers found.
H
Hangrui Xu
Shenzhen International Graduate School, Tsinghua University
Zhengxian Wu
Zhengxian Wu
Tsinghua University
Computer Vision、Large Language Model
Y
Yunyao Yu
Shenzhen International Graduate School, Tsinghua University
Z
Zhuohong Chen
Shenzhen International Graduate School, Tsinghua University
R
Rui Cong
Shenzhen International Graduate School, Tsinghua University
X
Xiangwen Deng
University of Arizona
Zhifang Liu
Zhifang Liu
School of Mathematical Sciences, Tianjin Normal University
image processing
P
Peng Jiao
Shenzhen International Graduate School, Tsinghua University
H
Haoqian Wang
Shenzhen International Graduate School, Tsinghua University