Pixel-Space Diffusion via Observation Operators

๐Ÿ“… 2026-08-22
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
ๆœฌๆ–‡้’ˆๅฏนๅƒ็ด ็ฉบ้—ดๆ‰ฉๆ•ฃๆจกๅž‹ไผ˜ๅŒ–้šพ้ข˜๏ผŒๆๅ‡บ่ง‚ๅฏŸ็ฎ—ๅญๆ‰ฉๆ•ฃๆก†ๆžถ๏ผŒ้€š่ฟ‡ไปŽ็ฒ—ๅˆฐ็ป†็š„็›‘็ฃ่ฝจ่ฟนไธŽ็‰นๅพ็ป†ๅŒ–๏ผŒๆ้ซ˜็”Ÿๆˆ่ดจ้‡ๅนถๅŠ ้€Ÿๆ”ถๆ•›ใ€‚
๐Ÿ“ Abstract
Pixel-space diffusion models directly model image distributions but remain difficult to optimize. Recent methods alleviate this challenge through target reparameterization, while still relying on a fixed clean-image target throughout denoising. Through empirical analysis, we identify a scale-time mismatch: image structures become predictable from coarse to fine as noise decreases, whereas existing models are forced to predict the full image even under high noise, resulting in low-SNR gradients that hinder optimization. To resolve this mismatch, we propose Observation Operator Diffusion, a unified framework that aligns both the supervision trajectory and feature refinement with the intrinsic recovery order of image structures. Specifically, we replace fixed full-image supervision along the standard flow path with a time-indexed observation trajectory that evolves from coarse structures to the full image during denoising. This trajectory is instantiated with a family of Gaussian-Lanczos operators at varying observation scales, yielding a path-consistent training objective. We further introduce GL-CoDA, a decoder that injects scale-specific Gaussian-Lanczos observations across decoding stages for coarse-to-fine feature refinement. Extensive experiments show that the proposed approach converges substantially faster while consistently improving generation quality, achieving an FID of 1.52 on ImageNet-256.
Problem

Research questions and friction points this paper is trying to address.

Pixel-space diffusion
Optimization
Denoising
Low-SNR gradients
Image structures
Innovation

Methods, ideas, or system contributions that make the work stand out.

Observation Operator Diffusion
scale-time mismatch
Gaussian-Lanczos operators
coarse-to-fine feature refinement
GL-CoDA
S
Shaojie Guo
East China Normal University
L
Lichen Ma
JD.com
H
Haoyang Tong
JD.com
Y
Yu He
JD.com
Z
Zipeng Guo
JD.com
X
Xiaoan Liu
JD.com
F
Feng Yan
JD.com
Y
Yu Guo
JD.com
F
Fei Wang
JD.com
Junshi Huang
Junshi Huang
Meituan
Computer VisionNLPMachine Learning
Yan Wang
Yan Wang
Professor in East China Normal University
computer visionmedical image analysis