Towards Uniformity and Alignment for Multimodal Representation Learning
This work addresses the modality gap in multimodal representation learning induced by the InfoNCE objective, which manifests as a conflict between inter-modal alignment and uniformity, as well as intra-modal alignment inconsistencies. The paper proposes the first framework that decouples alignment and uniformity in multimodal learning, employing Hölder divergence–based alignment optimization alongside a dedicated uniformity loss to effectively mitigate these conflicts. Theoretically, the proposed objective is shown to serve as a valid proxy for the global Hölder divergence between multimodal distributions. Notably, the method requires no task-specific components and consistently improves performance across both discriminative tasks (e.g., retrieval) and generative tasks (e.g., UnCLIP), demonstrating its generality and effectiveness.