Grounding Free-Form Instructions for Fashion Complementary Image Generation

📅 2026-08-24
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过引入自由形式指令和多模态语言接地模型StyleFlow,解决了时尚互补图像生成中用户自然查询理解和语言具体性问题。
📝 Abstract
Fashion complementary image generation (CIG) aims to create garments that stylistically match a seed item based on user intent, making it a natural multimodal grounding problem where models must interpret language in visual context. Existing CIG benchmarks rely on rigid template prompts (e.g., "a photo of a skirt"), failing to reflect natural user queries and obscuring model behavior across levels of linguistic specificity. We introduce fashion complementary image generation with free-form instructions, a multimodal language-grounding setting where a model generates a compatible garment from a seed image and a natural-language instruction. To this end, we enrich three CIG benchmarks with low-, medium-, and high-specificity instructions generated by a vision-language model and validated by human annotators. We instantiate the task with StyleFlow, a Rectified Flow Matching model that jointly conditions on the seed image and instruction within a single multimodal transformer. Across image quality metrics, catalog-alignment analysis, ablations, and human evaluation, StyleFlow consistently produces instruction-aligned and stylistically coherent garments while reducing architectural complexity and inference cost relative to auxiliary-module approaches.
Problem

Research questions and friction points this paper is trying to address.

Fashion Complementary Image Generation
Free-Form Instructions
Multimodal Grounding
Innovation

Methods, ideas, or system contributions that make the work stand out.

free-form instructions
multimodal language-grounding
StyleFlow
Rectified Flow Matching
🔎 Similar Papers