There and Back Again: Bidirectional Diffusion Bridges for Multimodality Translation

📅 2026-08-27
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决多模态翻译的灵活性和双向性问题,提出BIT方法,通过随机微积分实现从文本到图像及反向的生成路径。
📝 Abstract
Multimodality translation (e.g., text-to-image) is a core generative AI task. However, existing approaches (1) follow generative paths that do not directly represent the source modality, limiting the flexibility of some sampling algorithms; and (2) are unidirectional, preventing inversion (e.g., image-to-text). We propose BIT: Bidirectional Image-Text Diffusion Bridges. In contrast to previous approaches, BIT starts directly from text and interpolates into images, providing (1) a source-aware generative path that enables diverse and flexible sampling algorithms; and (2) an endpoint-conditioned process that can be traversed from image to text, providing a unified, bidirectional generative framework. BIT is derived through stochastic calculus, yielding SDE forms amenable to simulation and tractable loss functions that scale to high dimensions. Our experiments show that BIT is competitive with denoising-diffusion and deterministic-flow baselines, and outperforms them on several vision--language and natural-science evaluations.
Problem

Research questions and friction points this paper is trying to address.

Multimodality Translation
Generative Path
Unidirectional
Source Modality
Inversion
Innovation

Methods, ideas, or system contributions that make the work stand out.

Bidirectional Diffusion
Multimodality Translation
Source-Aware Generative Path
Endpoint-Conditioned Process