Partition the Support, Reconstruct the Residual: Training-Free Sparse Attention for Video Generation and World Models

📅 2026-08-18
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决训练免费的稀疏注意力在视频生成中的应用问题,提出SparsePR方法,结合响应耦合分区和探针拟合残差重构,减少注意力重建误差,提高速度。
📝 Abstract
Training-free block-sparse attention can accelerate video transformers, but row-wise attention concentration does not by itself specify an executable sparse operator. Queries sharing a block route may have poorly overlapping supports, while retained attention mass alone does not determine the post-softmax error from skipped interactions. We show that partition geometry affects both pooled support and the predictability of the remaining residual from the sparse output. We introduce SparsePR, which combines Response-Coupled Partitioning with Probe-Fitted Residual Reconstruction. Sampled-query key responses form paired K/V groups, whose centroids induce query-response coordinates for shared routing. A small set of exact query rows then calibrates a call-specific affine correction from the sparse output within the output subspace observed in the probe residuals. Across four heterogeneous video generation and world models, SparsePR consistently reduces attention-reconstruction error. Ablations show that probe fitting accounts for most of this reduction, while response-coupled partitioning lowers hard-drop error and improves reconstruction under a finite probe budget. SparsePR preserves generation quality at 22.0-26.0% realized executed-pair density while achieving 1.48x-2.61x end-to-end speedups. Project page: https://pardistaghavi.github.io/SparsePR-website/
Problem

Research questions and friction points this paper is trying to address.

Training-free
Sparse Attention
Video Generation
World Models
Attention-reconstruction Error
Innovation

Methods, ideas, or system contributions that make the work stand out.

SparsePR
Response-Coupled Partitioning
Probe-Fitted Residual Reconstruction
attention-reconstruction error