Fine-grained Textual Inversion Network for Zero-Shot Composed Image Retrieval
In zero-shot compositional image retrieval (ZS-CIR), single-pseudo-word mapping fails to capture fine-grained image semantics. To address this, we propose a fine-grained text inversion framework that decomposes an image into multiple pseudo-words—separately encoding subject and attribute semantics—and applies semantic regularization using BLIP-generated triplet-style captions. We further introduce multi-pseudo-word embedding modeling, template-guided cross-modal alignment, and contrastive learning to enhance compositional reasoning. Crucially, our method requires no annotated triplets. Evaluated on FashionIQ, CIRR, and CIRCO, it achieves significant improvements over state-of-the-art ZS-CIR approaches. These results demonstrate that fine-grained pseudo-word representations are essential for effective vision–language co-understanding in compositional retrieval tasks.