Sekai2: From World Exploration to Interactive World Modeling

📅 2026-08-10
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Existing video datasets struggle to simultaneously provide long-duration, multi-view real-world videos with aligned camera trajectories and temporally synchronized semantic annotations, thereby hindering the development of long-horizon interactive world models. To address this limitation, this work introduces a large-scale real-world video dataset that, for the first time, integrates complete camera trajectories, fine-grained hierarchical semantic annotations—encompassing subject motion, environmental dynamics, static content, and camera behavior—and nonlinear panoramic sequences supporting loops and revisits. The dataset comprises 2,826 hours of video (128,892 clips), 649,597 temporally aligned semantic segments, and 982 revisitable panoramic sequences spanning 113 countries, featuring full pose coverage, subtitles, and high semantic diversity, significantly advancing long-term spatial memory modeling and interactive generative capabilities.
📝 Abstract
Video world models must capture how scenes evolve over time and across viewpoints. Training them for long-horizon generation and camera control therefore benefits from long videos paired with camera trajectories and temporally grounded semantics. Existing corpora rarely offer the three together: large-scale web video provides broad visual diversity but no trajectories or time-aligned text, while pose-annotated datasets are typically short-range or reconstruction-oriented. We introduce Sekai2, a multi-source real-world video dataset that carries the world-exploration footage of Sekai toward interactive world modeling. The release contains 128,892 clips totaling 2,826 hours from 10,428 source videos across 113 countries or regions, and is deliberately weighted toward sustained observation: under a common 120-second decomposition, 43,594 segments reach the full two minutes and account for 51.4% of all footage. Every clip includes a released camera trajectory and hierarchical annotations disentangling subject motion, environment dynamics, static scene content, and camera behavior, resulting in 649,597 temporally grounded segments. Crucially, we further introduce 982 panoramic sequences captured along non-linear trajectories with loops and revisits. These revisits provide repeated observations of the same locations across time and viewpoints, offering essential supervision for learning persistent scene representations, long-term spatial memory, and geometrically consistent world models. Corpus-scale analyses demonstrate complete pose-and-caption coverage, broad geographic and semantic diversity, varied camera trajectories, and highly non-redundant temporal descriptions. Together, these properties make Sekai2 a scalable resource for long-horizon video generation, camera-controllable synthesis, and interactive world-model pre-training.
Problem

Research questions and friction points this paper is trying to address.

video world models
long-horizon generation
camera trajectories
temporally grounded semantics
interactive world modeling
Innovation

Methods, ideas, or system contributions that make the work stand out.

interactive world modeling
camera trajectory
temporally grounded semantics
panoramic revisits
long-horizon video generation
🔎 Similar Papers
No similar papers found.