Discrete Diffusion Bridges for Spatiotemporally Aligned Image Translation and Generation

📅 2026-08-30
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出离散扩散桥(DDB)框架,通过混合吸收机制和信息引导噪声调度解决图像转换和生成中的时空错位问题。
📝 Abstract
We propose Discrete Diffusion Bridges (DDB), a novel framework designed to resolve the fundamental spatiotemporal misalignment of standard discrete diffusion in image translation and generation. By corrupting data into a pure mask state via a random schedule, the conventional forward process induces a twofold misalignment: spatially, this pure-mask destination entirely discards the rich structural priors of the source image; temporally, the random masking order inherently contradicts the ``easy-first, hard-last'' decoding mechanism used during inference. To address this, DDB constructs a direct and efficient trajectory between domains. Spatially, we introduce a hybrid absorption mechanism that redefines the absorbing state to a stochastic mixture of mask and source tokens, effectively injecting source prior as spatial anchors into the latent space. Temporally, we design an information-guided noise schedule that quantifies semantic variation to prioritize the corruption of high-information regions at earlier timesteps. This ensures the model learns to resolve difficult semantic changes using robust context from invariant regions. Extensive experiments validate the versatility and robustness of our framework across diverse generative paradigms. DDB effectively balances edit alignment with structural fidelity across both text-guided semantic manipulation and pure structural image translation, while inherently complementing text-to-image generation and guaranteeing robust high-quality decoding under extremely low sampling steps. Code and models are available at \href{https://github.com/HKU-HealthAI/DDB}{https://github.com/HKU-HealthAI/DDB}.
Problem

Research questions and friction points this paper is trying to address.

Discrete Diffusion
Spatiotemporal Misalignment
Image Translation
Image Generation
Mask State
Innovation

Methods, ideas, or system contributions that make the work stand out.

Discrete Diffusion Bridges
spatiotemporal misalignment
hybrid absorption mechanism
information-guided noise schedule
🔎 Similar Papers
2024-03-19International Conference on Learning RepresentationsCitations: 4
💼 Related Jobs
No related jobs found.
X
Xing Xie
State Key Laboratory of Robotics and Intelligent Systems, Shenyang Institute of Automation, Chinese Academy of Sciences, Shenyang, China
Jiawei Liu
Jiawei Liu
Research Assistant Professor, Shenyang Institute of Automation, Chinese Academy of Sciences
computer visiondeep learningmultimediarobotics
S
Shijun Zhou
State Key Laboratory of Robotics and Intelligent Systems, Shenyang Institute of Automation, Chinese Academy of Sciences, Shenyang, China; University of Chinese Academy of Sciences, Beijing, China
Huijie Fan
Huijie Fan
Shenyang Institute of Automation, Chinese Academy of Sciences
Zhi Han
Zhi Han
SIA, CAS
Computer Vision
Yandong Tang
Yandong Tang
中国科学院沈阳自动化研究所教授
计算机视觉、图像处理、模式识别
Liangqiong Qu
Liangqiong Qu
The University of Hong Kong
Medical Image AnalysisImage SynthesisIllumination ModelingFederated Learning