Zing-0.5: Toward Playable Worlds with Real-Time Joint Action and Text Control

📅 2026-09-15
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出Zing-0.5模型,通过统一动作与文本条件、事件级监督及低成本实时交互方法,解决生成可游玩世界的实时联合控制问题。
📝 Abstract
We introduce Zing-0.5, a 5B autoregressive world model designed for playability: users can explore generated worlds, influence unfolding events, and respond to the resulting feedback through joint keyboard and online text control. Our approach brings together three technical contributions: (1) Unified action and text conditioning, combining magnitude-aware keyboard inputs with temporally aligned text instructions and jointly annotated videos to learn navigation and event control within the same sequence; (2) Event-scale supervision for incremental generation, using a segment-level teacher trained on connected multi-prompt videos to supervise a block-level causal student through distribution-matching distillation; and (3) Low-cost real-time interaction, combining four-step generation with context-preserving streaming to support 832 x 480 inference at 24 FPS at an estimated server rental cost of approximately USD 0.009 per stream-minute. Zing-0.5 achieves an overall score of 81.0 and a consistency score of 88.5 across 158 WBench Navigation cases. A joint-control demonstration shows a text-directed event change during continued navigation without restarting generation. We release the model weights, inference code, and Zing-SGLang serving implementation to support further work on playable generated worlds.
Problem

Research questions and friction points this paper is trying to address.

playable worlds
real-time joint action
text control
autoregressive world model
Innovation

Methods, ideas, or system contributions that make the work stand out.

Unified action and text conditioning
Event-scale supervision
Low-cost real-time interaction
🔎 Similar Papers
No similar papers found.