Reason What Matters: Retrieval-Grounded Reasoning for Universal Multimodal Embeddings

📅 2026-09-14
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决现有方法在大规模部署中的局限,提出ReWAM框架,通过检索反馈指导信用分配和推理计算,提高多模态嵌入的检索性能与效率。
📝 Abstract
Universal multimodal embedding (UME) learns unified representations across modalities, enabling a single model to support diverse retrieval tasks. Recent methods use Chain-of-Thought (CoT) reasoning to better interpret multimodal inputs before generating embeddings for complex retrieval tasks and further optimize this reasoning process through GRPO with retrieval-based rewards. However, two limitations hinder corpus-scale deployment. GRPO assigns all CoT tokens the same advantage, without identifying input-supported claims or evidence that distinguishes the positive from negatives. Moreover, generating a complete CoT before each embedding introduces substantial latency, even when a partial trace already provides sufficient retrieval evidence. To address these limitations, we propose Reason What Matters (ReWAM), a retrieval-grounded reasoning framework that uses retrieval feedback to guide both credit assignment and reasoning computation. Specifically, we introduce Retrieval-aware Self-Distillation (RASD), which constructs privileged guidance from input-supported evidence that distinguishes the positive item from retrieved hard negatives. An on-policy self-teacher uses this guidance to refine trajectory-level feedback into token-specific supervision for retrieval-relevant reasoning. We further develop Retrieval-adaptive Inference (RAI), which uses a retrieval confidence head to estimate the remaining retrieval utility of a partial CoT. It stops unproductive traces early and accelerates useful continuations with speculative decoding. Extensive experiments on MMEB-V2 and MRMR demonstrate that ReWAM achieves state-of-the-art retrieval performance while delivering up to 5x the inference throughput of competitive explicit-CoT UME methods. These results bridge the gap between retrieval quality and inference efficiency, making reasoning-enhanced UME practical for large-scale deployment.
Problem

Research questions and friction points this paper is trying to address.

Universal Multimodal Embedding
Chain-of-Thought Reasoning
Retrieval Efficiency
Innovation

Methods, ideas, or system contributions that make the work stand out.

Retrieval-Grounded Reasoning
Retrieval-aware Self-Distillation
Retrieval-adaptive Inference
Universal Multimodal Embeddings
🔎 Similar Papers
No similar papers found.
M
Mingzhou Jiang
Tsinghua Shenzhen International Graduate School, Tsinghua University
Peixi Wu
Peixi Wu
University of Science and Technology of China
MultiModalNeuromorphicObject Detection
H
Hang Cheng
Tsinghua Shenzhen International Graduate School, Tsinghua University
Yunhao Zhou
Yunhao Zhou
Shanghai Jiao Tong University
EDAGNNLLM
Biao Yang
Biao Yang
Shanghai Jiao Tong University, Antai College of Economics and Management
Asset PricingClimate Finance
W
Wei Yuan
Kuaishou Technology
Y
Yun Li
College of Future Information Technology, Fudan University
F
Fan Yang
Kuaishou Technology
W
Wenwu Ou
Kuaishou Technology
H
Honghui He
Tsinghua Shenzhen International Graduate School, Tsinghua University