Factorized Transport Alignment for Multimodal and Multiview E-commerce Representation Learning
Existing vision-language models (VLMs) align only image titles with primary images, neglecting the critical semantic information conveyed by non-primary images and auxiliary textual modalities (e.g., descriptions, tags) in e-commerce scenarios—thereby limiting multimodal, multi-view representation learning. To address this, we propose Factorized Transport: a lightweight, factorized optimal transport approximation method designed for open-platform e-commerce. It enables scalable multi-view alignment—spanning primary/auxiliary images and titles/descriptions/tags—while supporting zero-overhead online inference fusion. Our approach integrates stochastic view sampling, dual-tower embedding caching, and multi-view contrastive learning. Evaluated on a million-scale industrial product dataset, it achieves a +7.9% improvement in Recall@500 over strong multimodal baselines, demonstrating both effectiveness and deployability for large-scale, real-time e-commerce search.