Your Embedding Model is SMARTer Than You Think
This work addresses the limitations of unimodal single-vector retrieval models, which struggle to preserve fine-grained local information in multimodal tasks, and existing multi-vector approaches that typically require retraining and lack effective global representations. To overcome these challenges, the authors propose SMART, a plug-and-play framework that activates the latent multi-vector capabilities embedded in frozen single-vector models during inference—without any additional training. By fusing global and local information through late interaction of intermediate hidden states, SMART achieves both efficient inference and lightweight adaptation. Experimental results demonstrate that SMART significantly outperforms state-of-the-art multi-vector models on the MMEB-V2 benchmark and visual document retrieval tasks, delivering superior retrieval performance at lower computational cost.