PlayWorld: Benchmarking World Models with Agent Players over Long-Horizon Objectives

📅 2026-08-13
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Existing video world models lack a unified evaluation framework tailored to long-horizon, goal-driven interactive tasks, making it difficult to assess their consistency and controllability. To address this gap, this work proposes the first goal-oriented, multimodal agent interaction benchmark, comprising 171 diverse scenarios, and introduces a systematic evaluation protocol across four key dimensions: geometric consistency, interaction fidelity, out-of-view dynamics, and insight evolution. Empirical results reveal that nine state-of-the-art models exhibit significant shortcomings in spatial coherence and persistent state evolution, undermining their reliability in complex interactive settings. This study establishes a new standard and methodological foundation for comprehensive evaluation of world models.
📝 Abstract
Video world models simulate future states conditioned on current observations and user actions. Recent systems have demonstrated impressive video consistency and action controllability over long sequences. However, fairly comparing these interactive models remains challenging. In practice, a human player typically evaluates a world model by pursuing long-horizon objectives through interaction. For example, a user may turn around 360 degrees to see whether the environment remains consistent, or walk into the water and inspect whether realistic water ripples are generated. The action sequence required to achieve the same objective may vary substantially between models, making fixed action-conditioned evaluation unsuitable for cross-model comparison. To address this, we employ multi-modal Agent Players to interact with world models toward specified long-horizon objectives. Building on this paradigm, we introduce PlayWorld, a benchmark providing 171 scenarios, each with a specified objective. To evaluate performance thoroughly, we assess models along four core dimensions: geometry consistency, interaction fidelity, out-of-sight evolution, and insight evolution. In addition, we incorporate basic ability metrics for video quality and controllability. Experiments across nine state-of-the-art world models reveal that current models remain unreliable on long-horizon interactive objectives, particularly in maintaining spatial consistency and persistent state evolution. Code and data are available at https://github.com/kxding/PlayWorld.
Problem

Research questions and friction points this paper is trying to address.

world models
long-horizon objectives
interactive evaluation
agent players
benchmarking
Innovation

Methods, ideas, or system contributions that make the work stand out.

world models
agent players
long-horizon objectives
interactive evaluation
spatial consistency
🔎 Similar Papers
No similar papers found.