Accelerating Diffusion Transformers with Gaussian Process Rectified Feature Cache

📅 2026-09-05
📈 Citations: 0
Influential: 0
📄 PDF
📝 Abstract
Diffusion Transformers have become the dominant paradigm in generative AI, but their high computational costs severely hinder real-time applications. Prediction-based feature caching is widely used to accelerate diffusion transformers; however, as the number of steps increases, the deviation between its predictions and the reference full-compute trajectory gradually grows. An intuitive idea is to use an online regression model to dynamically correct this deviation, but it faces the issue of label data being unavailable during the acceleration process. This paper presents a statistical observation that the residuals between the features of full computation steps using caching methods and reference full-compute trajectory locally exhibit a zero-mean Gaussian distribution. By treating the features of full computation steps as noisy observations of reference features, the data acquisition problem is resolved. Based on this observation, a plug-and-play GP-Refiner correction framework is proposed. This method utilizes Gaussian Process Regression for correction and, leveraging the properties of GPR, introduces an uncertainty-adaptive computation strategy that triggers necessary full-computation calibration by monitoring the posterior variance in real time. Experiments demonstrate significant improvements across different models when combined with various state-of-the-art methods. Integrating the proposed framework with TaylorSeer reduces the computational load by 19.3% while improving PSNR by 0.9 dB and reducing LPIPS from 0.46 to 0.29. Code is available in https://github.com/Aredstone/GP-Refiner.
Problem

Research questions and friction points this paper is trying to address.

Diffusion Transformers
computational costs
real-time applications
feature caching
deviation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Gaussian Process Regression
Feature Cache
Diffusion Transformers
Adaptive Computation
Posterior Variance
🔎 Similar Papers
Z
Zhirong Shen
Shanghai Jiao Tong University
R
Rui Huang
Shanghai Jiao Tong University
Chang Zou
Chang Zou
Intern at EPIC Lab, Shanghai Jiao Tong University
Generative modelsImages and Videos generation
S
Shikang Zheng
Shanghai Jiao Tong University
J
Jiacheng Liu
Shanghai Jiao Tong University
P
Peiliang Cai
Shanghai Jiao Tong University
Z
Zhengyi Shi
Xiamen University
Yaosong Du
Yaosong Du
UESTC, DAIL Tech
Advanced Routing AlgorithmReinforcement LearningCVMedical Agent SystemLLM
L
Liang Feng
Fudan University
X
Xiaobing Tu
Terminal Intelligent Computing Division, Alibaba Cloud
J
Jinkui Ren
Terminal Intelligent Computing Division, Alibaba Cloud
Xiantao Zhang
Xiantao Zhang
Beihang University
Large Language ModelsNatural Language ProcessingArtificial IntelligenceData Curation
Linfeng Zhang
Linfeng Zhang
DP Technology; AI for Science Institute
AI for Sciencemulti-scale modelingmolecular simulationdrug/materials design