π€ AI Summary
This work addresses the challenges of manifold drift and inadequate spatial structure representation in autoregressive prediction for continuous environment modeling. To overcome these issues, the authors propose a self-stabilizing, action-conditioned Joint Embedding Predictive Architecture (JEPA) that maps the environment onto a discrete two-dimensional grid. By incorporating translation-equivariant constraints and a grid βsnappingβ mechanism, the model achieves structured world representation. A topological quantization strategy based on balanced continuous entropy regularization is introduced to embed error correction directly into the latent space, thereby mitigating manifold drift. Furthermore, maximum-entropy exploration encourages learning of generalizable dynamics rather than memorizing trajectories. Experiments demonstrate that the model functions effectively as both a spatial physics simulator and a causal discovery system across passive observation, active control, and abstract sequential tasks, successfully constructing structured topological maps.
π Abstract
We present the Global Neural World Model (GNWM), a self-stabilizing framework that achieves topological quantization through balanced continuous entropy constraints. Operating as a continuous, action-conditioned Joint-Embedding Predictive Architecture (JEPA), the GNWM maps environments onto a discrete 2D grid, enforcing translational equivariance without pixel-level reconstruction. Our results show this architecture prevents manifold drift during autoregressive rollouts by using grid ``snapping'' as a native error-correction mechanism. Furthermore, by training via maximum entropy exploration (random walks), the model learns generalized transition dynamics rather than memorizing specific expert trajectories. We validate the GNWM across passive observation, active agent control, and abstract sequence regimes, demonstrating its capacity to act not just as a spatial physics simulator, but as a causal discovery model capable of organizing continuous, predictable concepts into structured topological maps.