SAGE: Semantic Attribute Graphs for Multi-Entity Visual Retrieval

📅 2026-09-01
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出SAGE框架,通过构建语义属性图解决密集文档图像中实体信号混淆问题,提高细粒度视觉检索精度。
📝 Abstract
Dense document images often contain many fine-grained visual and textual entities whose relevance depends on a user query. Standard vision-language retrievers encode cropped regions with a single vector, which can mix distinct entity signals and obscure the evidence needed for fine-grained retrieval. We call this failure mode Semantic Dilution and quantitatively show that it degrades entity-level retrieval as a function of entity density. To mitigate it, we propose SAGE, a training-free framework that parses semantic entities from dense document images, represents them as hierarchical graph nodes with multi-vector embeddings, and retrieves query-relevant evidence through iterative entity-level subgraph matching. We also introduce DEAR, a dataset of 1,055 query--image pairs sourced from product detail pages, where each query requires retrieving and comparing multiple fine-grained entities from visually dense inputs across four question types of increasing complexity. Experiments show that SAGE substantially reduces semantic dilution and outperforms patch-level and OCR-based retrieval baselines on DEAR, achieving a Recall@3 of 0.849 and a generation score of 2.746 on multi-entity visual comparison queries. Our code is available at https://github.com/All4Nothing/SAGE.
Problem

Research questions and friction points this paper is trying to address.

Semantic Dilution
multi-entity visual retrieval
dense document images
Innovation

Methods, ideas, or system contributions that make the work stand out.

Semantic Attribute Graphs
multi-vector embeddings
entity-level subgraph matching
Semantic Dilution
🔎 Similar Papers
No similar papers found.
Y
Yongjoo Kim
Korea University
M
Mincheol Kwon
Korea University
S
Seonga Choi
Korea University
M
Minseung Lee
Korea University
K
Kyeong-Jin Oh
KT Corporation
H
Hyunyoung Lee
KT Corporation
Y
Yunsu Choi
KT Corporation
Jungbeom Lee
Jungbeom Lee
Amazon
Deep LearningComputer VisionMulti-Modal Learning