Institution profile

Harbin Institute of Technology

Academic institutionasia · cn
Official website
Research library2,420linked papers
Opportunities0open roles
Selected work

Representative Papers

COutfitGAN: Learning to Synthesize Compatible Outfits Supervised by Silhouette Masks and Fashion Styles

Feb 12, 2025IEEE transactions on multimedia

This paper introduces a novel task—fashion outfit generation from an arbitrary number of given clothing items—aiming to synthesize visually realistic and stylistically harmonious complementary garments. Methodologically, it proposes the first generative framework for completing outfits from partial inputs, featuring a pyramid-style extractor to model multi-granularity fashion features, and a dual-discriminator joint optimization scheme: a U-Net-based discriminator assesses image realism, while a relational discriminator models cross-item compatibility. Additionally, contour mask supervision is incorporated to enhance fine-grained structural consistency. Evaluated on a large-scale dataset comprising 200K outfits and 800K individual items, the method achieves significant improvements over state-of-the-art approaches across quantitative metrics—including image fidelity, outfit compatibility, and visual similarity—demonstrating both effectiveness and generalizability.

23 citationsRead paper

Dynamic Multimodal Fusion via Meta-Learning Towards Micro-Video Recommendation

Aug 30, 2023ACM Trans. Inf. Syst.

To address the limitation of static multimodal fusion in middle-school micro-video recommendation—its inability to capture inter-video modality relationship discrepancies—this paper proposes MetaMMF, a meta-learning-based dynamic multimodal fusion framework. Methodologically, MetaMMF treats multimodal fusion for each video as an individual meta-task and employs meta-learning to generate video-specific fusion functions; it further adopts CP tensor decomposition to enhance parameter efficiency and training stability. While implicitly incorporating graph neural network principles (e.g., akin to MMGCN), MetaMMF avoids explicit graph construction. Extensive experiments on three benchmark datasets demonstrate that MetaMMF consistently outperforms state-of-the-art models—including MMGCN, LATTICE, and InvRL—achieving superior recommendation accuracy and computational efficiency. The source code is publicly released, empirically validating the dual advantages of dynamic fusion in both performance and efficiency.

15 citationsRead paper

Learning to Synthesize Compatible Fashion Items Using Semantic Alignment and Collocation Classification: An Outfit Generation Framework

Sep 15, 2022IEEE Transactions on Neural Networks and Learning Systems

This work addresses the challenging problem of complete outfit generation conditioned on a single garment and target-region masks—a key task in fashion design automation. We propose OutfitGAN, an end-to-end generative framework that synthesizes compatible tops, bottoms, footwear, and accessories given an input garment and spatially localized masks. Methodologically, we introduce two novel components: (i) a Semantic Alignment Module (SAM) that models fine-grained cross-garment semantic correspondences, and (ii) a Compatibility Classification Module (CCM) that explicitly enforces style and semantic coherence. Our multi-stage GAN architecture integrates semantic segmentation guidance, feature-level alignment losses, compatibility-aware adversarial supervision, and mask-conditioned generation control. Evaluated on a large-scale dataset of 20,000 real-world outfits, OutfitGAN achieves state-of-the-art performance across image fidelity, perceptual realism, and outfit compatibility metrics. It enables high-fidelity, diverse, and interactive fashion editing.

13 citationsRead paper

Fine-grained Textual Inversion Network for Zero-Shot Composed Image Retrieval

Jul 10, 2024Annual International ACM SIGIR Conference on Research and Development in Information Retrieval

In zero-shot compositional image retrieval (ZS-CIR), single-pseudo-word mapping fails to capture fine-grained image semantics. To address this, we propose a fine-grained text inversion framework that decomposes an image into multiple pseudo-words—separately encoding subject and attribute semantics—and applies semantic regularization using BLIP-generated triplet-style captions. We further introduce multi-pseudo-word embedding modeling, template-guided cross-modal alignment, and contrastive learning to enhance compositional reasoning. Crucially, our method requires no annotated triplets. Evaluated on FashionIQ, CIRR, and CIRCO, it achieves significant improvements over state-of-the-art ZS-CIR approaches. These results demonstrate that fine-grained pseudo-word representations are essential for effective vision–language co-understanding in compositional retrieval tasks.

10 citations1 influentialRead paper

Towards Vision Enhancing LLMs: Empowering Multimodal Knowledge Storage and Sharing in LLMs

Nov 27, 2023arXiv.org

Existing multimodal large language models (MLLMs) employ only unidirectional mapping of visual information into the language space, failing to effectively leverage visual knowledge to enhance holistic reasoning. This work introduces Vision-Augmented Large Language Models (VA-LLMs), breaking from this paradigm by enabling LLMs to actively store, share, and retrieve visual knowledge. Our contributions are threefold: (1) Modular Visual Memory (MVM), a structured, long-term storage mechanism for visual knowledge; (2) Soft Multimodal Mixture of Experts (MoME), a dynamic architecture that coordinates vision and language experts during autoregressive generation; and (3) a cross-modal knowledge injection and collaborative reasoning framework. Experiments demonstrate substantial improvements in physical commonsense understanding and spatial reasoning, achieving state-of-the-art performance across multiple multimodal benchmarks, including MMMU, ScienceQA, and POPE.

6 citationsRead paper
Recent publications

Latest Papers