Attribute-Conditioned Multimodal Slot Factorization for Controllable Fashion Retrieval
This work addresses the limitation of conventional fashion retrieval methods, which conflate multi-attribute information into a single embedding, thereby hindering fine-grained control over specific attributes such as color or category. To overcome this, the authors propose the MM-slotgate model, which introduces— for the first time—a named and intervenable multimodal attribute-slot architecture. This framework decomposes Fashion-CLIP’s joint image-text embeddings into four semantically distinct attribute slots, each equipped with an independent image-text gating mechanism that automatically learns modality weights without explicit supervision. The approach achieves both semantic interpretability and controllable retrieval performance, attaining a ConstraintSatisfied@10 score of 0.7566 on the H&M dataset, improving color retrieval success by 0.568, and yielding a 15.3× enhancement in color intervention efficacy through quantized slot representations.