PailitaoGR: Latent Think-with-Images for Generative Image Retrieval

📅 2026-08-27
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出PailitaoGR方法,通过目标聚焦感知和选择性利用辅助证据解决生成式图像检索中的多样化信息处理问题。
📝 Abstract
Generative retrieval has demonstrated strong performance by directly generating product semantic identifiers (SIDs). Extending this paradigm to image search, however, is nontrivial because real-world query images contain diverse information, including the search target, useful auxiliary evidence, and irrelevant visual content. This requires the model to identify and focus on the search target while selectively utilizing auxiliary evidence. In this paper, we propose \textbf{PailitaoGR}, a \emph{Latent Think-with-Images} method for generative image retrieval, which internalizes target-focused perception and selective auxiliary-evidence utilization into a the generative retrieval model, enabling \textit{Zooming without Cropping} and \textit{Reading without OCR}. Specifically, we design a target-focused perception mechanism that identifies and enhances visual tokens of the search target, consisting of a target Enhancer and a learning strategy based on on-policy distillation and attention guidance loss, enabling the model to focus on search-target regions. We also design a selective auxiliary-evidence utilization mechanism that identifies and enhances visual tokens of auxiliary evidence, including an auxiliary enhancer and an in-capacity incremental contrastive distillation strategy, enabling the model to exploit auxiliary evidence. We construct training and validation sets sampled from real-world online image-search logs. Experiments show that our method outperforms existing baselines by an average of 13.8\%, validating its effectiveness.
Problem

Research questions and friction points this paper is trying to address.

generative retrieval
search target
auxiliary evidence
visual content
Innovation

Methods, ideas, or system contributions that make the work stand out.

Latent Think-with-Images
Generative Image Retrieval
Target-focused Perception
Selective Auxiliary-evidence Utilization
Xiaomeng Fan
Xiaomeng Fan
Beijing Institute of Technology
machine learningcomputer vision
Y
Yueran Liu
Alibaba Group
S
Shengyu Zhou
Alibaba Group
C
Chenghan Fu
Alibaba Group
W
Wanxian Guan
Alibaba Group
F
Feng Li
Alibaba Group
C
Chuan Yu
Alibaba Group
Jian Xu
Jian Xu
Senior Director, Ad Platform, Alibaba Group
Computational AdvertisingMachine LearningData MiningData Privacy
B
Bo Zheng
Alibaba Group