FashionLens: Toward Versatile Fashion Image Retrieval via Task-Adaptive Learning

📅 2026-05-21
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Existing fashion image retrieval methods are often confined to single-task settings, limiting their ability to accommodate diverse query modalities and user intents. This work proposes FashionLens, a unified framework built upon multimodal large language models that enables task-adaptive learning for versatile retrieval scenarios. The approach introduces two key innovations: a Proposal-Guided Spherical Query Calibrator that dynamically refines query representations, and a Gradient-Guided Adaptive Sampling strategy to balance optimization across multiple tasks. Evaluated on the newly curated U-FIRE benchmark, FashionLens substantially outperforms current state-of-the-art methods and demonstrates exceptional cross-task generalization. The code and dataset are publicly released to facilitate further research.
📝 Abstract
Fashion image retrieval is a cornerstone of modern e-commerce systems. A unified framework that supports diverse query formats and search intentions is highly desired in practice. However, existing approaches focus on narrow retrieval tasks and do not fully capture such diversity. Therefore, in this work, we aim to develop a unified framework capable of handling diverse realistic fashion retrieval scenarios, achieving truly versatile fashion image retrieval. To establish a data foundation, we first introduce U-FIRE, a comprehensive benchmark that consolidates fragmented fashion datasets into a unified collection, supplemented by two manually curated datasets for testing generalization. Building upon this, we propose FashionLens, a unified framework based on Multimodal Large Language Models. To handle divergent matching objectives, we design a Proposal-Guided Spherical Query Calibrator that dynamically shifts query representations into task-aligned metric spaces via adaptive spherical linear interpolation. Additionally, to mitigate the optimization imbalance caused by varying task complexities and data scales, we develop a Gradient-Guided Adaptive Sampling strategy that automatically re-weights tasks based on realtime learning difficulty and the data scale prior. Experiments on U-FIRE show that FashionLens achieves state-of-the-art performance across diverse retrieval scenarios and generalizes robustly to unseen tasks. The data and code are publicly released at https://github.com/haokunwen/FashionLens.
Problem

Research questions and friction points this paper is trying to address.

fashion image retrieval
unified framework
diverse query formats
search intentions
versatile retrieval
Innovation

Methods, ideas, or system contributions that make the work stand out.

Task-Adaptive Learning
Multimodal Large Language Models
Spherical Query Calibration
Gradient-Guided Adaptive Sampling
Unified Fashion Retrieval
🔎 Similar Papers
No similar papers found.
Haokun Wen
Haokun Wen
Harbin Institute of Technology, Shenzhen
Multimedia ComputingInformation Retrieval
Xuemeng Song
Xuemeng Song
City University of Hong Kong
Information RetrievalMultimedia Analysis
X
Xinghao Xie
School of Artificial Intelligence, Nanjing University, Nanjing 210023, China
X
Xiaolin Chen
Institute of Data Science, National University of Singapore, Singapore
Xiangyu Zhao
Xiangyu Zhao
Associate Professor, City University of Hong Kong
RecommendationsLarge Language Models (LLMs)TrustworthyAISearch EngineUrban Computing
W
Weili Guan
School of Information Science and Technology, Harbin Institute of Technology (Shenzhen), Shenzhen 518055, China; and also with the Shenzhen Loop Area Institute, Shenzhen 518045, China