WALL-SS: Scaling Long-horizon World Models via Next-Scale Autoregression

📅 2026-08-26
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
WALL-SS通过尺度自回归扩展生成长期世界模型,解决机器人模拟中动作可控性和长时预测的问题。
📝 Abstract
Generative world models provide robots with predictive models of how the world evolves under interaction, with growing potential for simulation, planning, policy evaluation, and robot learning. Beyond clip-level future prediction, a unified generative formulation should relate actions to consequences, support flexible horizons and continuous interaction, and enable reward-driven optimization. We introduce WALL-SS, a world model that generates visual futures through Scale-wise autoregressive Scaling, enabling action-controllable and long-horizon robotic simulation. WALL-SS represents embodied trajectories as causal sequences of temporally interleaved observations and actions, making action-dependent state transitions explicit while naturally supporting variable-length generation, streaming extension through reusable causal states, and direct optimization through sequence probabilities. To make this formulation effective over long horizons, we generate each future observation in a coarse-to-fine manner and develop three complementary components within the same hierarchy. Action-conditioned next-scale prediction injects scale-aligned action representations to improve action-future coupling and model both successful and failed behaviors. Scale-compressed long-horizon memory retains recent interactions at fine resolution while compressing distant observations and actions, with scale-wise dream forcing enhancing robustness to self-generated context. Finally, on-policy alignment optimizes autoregressive visual dynamics with action-following and long-term consistency rewards while preserving the pretrained visual distribution. Experiments show that WALL-SS improves action following and trajectory accuracy, supports coherent minute-long streaming rollout under bounded memory, and consistently benefits from on-policy alignment in reducing action drift and long-horizon inconsistency.
Problem

Research questions and friction points this paper is trying to address.

Generative world models
Long-horizon prediction
Action-controllable simulation
Flexible horizons
Reward-driven optimization
Innovation

Methods, ideas, or system contributions that make the work stand out.

Next-Scale Autoregression
Scale-wise autoregressive Scaling
on-policy alignment
long-horizon memory
🔎 Similar Papers
No similar papers found.
M
Maeve Zhang
R
Rain Sun
X
Xiang Wang
Cyril Zhang
Cyril Zhang
OpenAI
learningalgorithms
S
Shalfun Li
Meng Cao
Meng Cao
Postdoc, Carnegie Mellon University
Psychology
H
Howard Lu
Ethan Chen
Ethan Chen
University of Rochester
Computer Science
H
Harry Jhou
K
KZ Zheng
L
Lights Shi
R
Regis Cheng
L
Lorenzin
R
Robert Wang
V
Victor Yao
G
Gody Li
E
Elise Mon
Y
Yohann Tang
R
Ryan Yu
P
PS Zhang
V
Vincent Chen
H
Hang Su
R
Roy Gan
H
Hao Wang
Q
Qian Wang