WorldRover: A Scalable Synthetic Video Data Engine for World Exploration with Rich Annotations

📅 2026-08-16
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the scarcity of long-horizon annotated data integrating geometric, kinematic, and control signals for world exploration models. We propose a scalable synthetic data engine built upon Unreal Engine that leverages offline rendering, synchronized multi-view acquisition, and trajectory-derived action signals to automate the generation of exploration videos featuring dense optical flow and 3D point tracking. Based on this framework, we introduce WorldRover-10M, a dataset comprising minute-level RGB imagery, metric depth, camera trajectories, and action annotations. By providing high-quality, scalable supervision, this work effectively bridges the critical data gap in world model training, enabling robust learning from long-sequence multimodal signals essential for embodied AI research.
📝 Abstract
Learning to generate or reconstruct explorable worlds requires video paired with more than RGB: camera motion, scene geometry, temporal correspondence and, for interactive models, control signals. Real capture can provide some of these signals, but dense geometry and long-range correspondence usually rely on estimation or specialised instrumentation. Rendering provides these quantities directly, yet existing synthetic resources rarely combine them on the same frames while also supporting controlled changes of viewpoint and appearance. We introduce WorldRover, a data engine for generating richly annotated, long-range explorations of artist-built environments. At its core, WorldRover-Engine is an Unreal Engine pipeline that executes and offline-renders minute-scale routes while preserving their full trajectories and scene geometry. The same exploration can be replayed from first-person, third-person, and 360 panoramic cameras under different environmental states. Using WorldRover-Engine, we construct WorldRover-10M, whose sequences pair RGB with metric depth, camera trajectories, and trajectory-derived action signals throughout each exploration. Third-person subsets additionally provide dense optical flow, long-range 2D/3D point tracks with visibility, and a character trajectory distinct from the camera trajectory. The engine can render a traversal from first-person, third-person and 360 panoramic viewpoints, under different environmental states or with a neutral white material, while preserving the route and scene geometry. WorldRover therefore turns long-horizon world exploration into a scalable data-generation problem, providing supervision for models that must build, maintain, and revisit coherent representations of an explorable world.
Problem

Research questions and friction points this paper is trying to address.

Synthetic Video Data
World Exploration
Rich Annotations
Scene Geometry
Scalable Data Generation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Synthetic Video Data Engine
Rich Annotations
Multi-view Rendering
Long-range Exploration
WorldRover-10M
🔎 Similar Papers
No similar papers found.