Beyond Fixed Formulas: Data-Driven Linear Predictor for Efficient Diffusion Models

📅 2026-04-29
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the high sampling cost of Diffusion Transformers (DiTs) and the failure of existing handcrafted feature caching methods under aggressive step skipping. To this end, the authors propose L2P (Learnable Linear Predictor), a data-driven feature caching framework that introduces, for the first time, a learnable per-timestep linear predictor into the caching mechanism. This lightweight module reconstructs current features from historical feature trajectories, thereby overcoming the limitations of fixed-formula approaches. Requiring only brief single-GPU training, L2P achieves a 4.55× reduction in FLOPs and a 4.15× speedup in inference latency on FLUX.1-dev, while enabling up to a 7.18× acceleration on the Qwen-Image model with minimal degradation in visual fidelity.
📝 Abstract
To address the high sampling cost of Diffusion Transformers (DiTs), feature caching offers a training-free acceleration method. However, existing methods rely on hand-crafted forecasting formulas that fail under aggressive skipping. We propose L2P (Learnable Linear Predictor), a simple data-driven caching framework that replaces fixed coefficients with learnable per-timestep weights. Rapidly trained in ~20 seconds on a single GPU, L2P accurately reconstructs current features from past trajectories. L2P significantly outperforms existing baselines: it achieves a 4.55x FLOPs reduction and 4.15x latency speedup on FLUX.1-dev, and maintains high visual fidelity under up to 7.18x acceleration on Qwen-Image models, where prior methods show noticeable quality degradation. Our results show learning linear predictors is highly effective for efficient DiT inference. Code is available at https://github.com/Aredstone/L2P-Cache.
Problem

Research questions and friction points this paper is trying to address.

Diffusion Transformers
feature caching
sampling acceleration
linear prediction
inference efficiency
Innovation

Methods, ideas, or system contributions that make the work stand out.

Learnable Linear Predictor
Diffusion Transformers
Feature Caching
Data-Driven Acceleration
Efficient Inference
🔎 Similar Papers
2024-04-19Neural Information Processing SystemsCitations: 14
Z
Zhirong Shen
Shanghai Jiao Tong University; University of Electronic Science and Technology of China
Rui Huang
Rui Huang
University of Electronic Science and Technology of China
exoskeletonreinforcement learningrobot vision
J
Jiacheng Liu
Shanghai Jiao Tong University; Shandong University
Chang Zou
Chang Zou
Intern at EPIC Lab, Shanghai Jiao Tong University
Generative modelsImages and Videos generation
P
Peiliang Cai
Shanghai Jiao Tong University
S
Shikang Zheng
Shanghai Jiao Tong University
Z
Zhengyi Shi
Xiamen University
L
Liang Feng
Fudan University
Linfeng Zhang
Linfeng Zhang
DP Technology; AI for Science Institute
AI for Sciencemulti-scale modelingmolecular simulationdrug/materials design