VideoTok4D: A 4D-Aware Video Tokenizer for Compact World Representation

📅 2026-09-11
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决视频tokenizer在4D场景紧凑表示上的局限,提出VideoTok4D,通过时空解耦、轨迹感知动态注意力机制和Co4DGen方法实现高效4D场景生成。
📝 Abstract
Video tokenizers have emerged as a cornerstone of modern video modeling, underpinning progress in compression, reconstruction and generation by mapping high-dimensional visual signals into compact latent spaces. However, despite this progress, current tokenization paradigms largely remain within the 2D visual domain, treating videos as image sequences rather than observations of an underlying dynamic 3D world. Consequently, the learned tokens inherit this observation-centric bias, limiting their capacity to compactly represent real-world 4D scenes. To mitigate this issue, we propose VideoTok4D, a novel 4D-aware video tokenizer for compact world representation. Specifically, our approach comprises three key designs: 1) a spatiotemporal disentanglement strategy that factorizes videos into static and dynamic tokens for holistic world modeling; 2) a track-aware dynamic attention mechanism that aggregates trajectory-aligned cues to promote cross-view motion consistency; and 3) Co4DGen, a diffusion prior learned over the resulting VideoTok4D token space for efficient 4D scene generation. Extensive experiments have demonstrated that our proposed method achieves state-of-the-art performance while requiring up to 4 orders of magnitude less storage than dense 4D representations. Moreover, the compact token space substantially shortens diffusion sequences, enabling efficient generation.
Problem

Research questions and friction points this paper is trying to address.

video tokenizers
4D-aware
dynamic 3D world
compact representation
spatiotemporal disentanglement
Innovation

Methods, ideas, or system contributions that make the work stand out.

4D-aware
spatiotemporal disentanglement
track-aware dynamic attention
Co4DGen
compact world representation
🔎 Similar Papers