MotionCraft: Latent World Modeling with Sparse Attention for Visual Upscaling

📅 2026-08-09
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenge in video super-resolution of simultaneously preserving fine local details, modeling long-range spatiotemporal dependencies, achieving perceptual realism, and maintaining computational efficiency. To this end, the authors propose a motion-aware latent state prediction framework featuring several key innovations: a motion-informed latent world representation, a Latent World Transformer that balances local and non-local interactions, an adaptive sparse attention mechanism, and a compact conditional decoder enabling user-controllable reconstruction. The proposed method achieves significant improvements in both reconstruction accuracy and perceptual quality while retaining high computational efficiency. Moreover, it offers a predictable trade-off between temporal smoothness and detail fidelity, allowing users to tailor the output according to specific application needs.
📝 Abstract
Video super-resolution (VSR) aims to recover high-fidelity high-resolution videos from low-resolution inputs and is central to applications ranging from mobile capture to streaming and archival restoration. Existing approaches trade off among local-detail fidelity, long-range spatio-temporal modeling, perceptual realism, and efficiency: convolutional alignment techniques preserve local structure but suffer when motion is large or degradations are complex; transformer-based methods capture long-range dependencies yet require architectural or algorithmic adaptations to remain computationally feasible; and recent latent or diffusion-based generators synthesize rich texture but require specialized temporal constraints to maintain coherence. We present MotionCraft, a controllable VSR framework that formulates restoration as motion-aware latent state prediction inspired by world models and integrates adaptive sparse attention with an explicit user-accessible control interface. MotionCraft combines robust motion fusion, a Latent World Transformer that balances locality and targeted non-local interactions, and a compact conditional decoder to deliver temporally consistent, high-quality reconstructions under streaming constraints. Empirical evaluations show that MotionCraft achieves strong reconstruction and perceptual performance while enabling predictable trade-offs between temporal smoothness and reconstruction fidelity.
Problem

Research questions and friction points this paper is trying to address.

video super-resolution
spatio-temporal modeling
perceptual realism
computational efficiency
temporal coherence
Innovation

Methods, ideas, or system contributions that make the work stand out.

latent world modeling
sparse attention
motion-aware super-resolution
controllable video restoration
temporal consistency
🔎 Similar Papers
No similar papers found.