ParallelWorld: Test-Time Scaling for Embodied Reasoning

📅 2026-08-24
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决复杂环境中延迟反馈问题,提出ParallelWorld框架,通过多步平行模拟和验证者引导的树搜索方法优化长期规划。
📝 Abstract
Embodied Reasoning constitutes a fundamental capability of embodied intelligence, serving as the basis for autonomous perception, reasoning, and interaction within physical environments. Recent studies have shifted the paradigm of embodied reasoning from static perception toward dynamic exploration, where agents acquire task-relevant information through interactions with the environment. However, existing active reasoning approaches generally generate exploration trajectories incrementally without long-horizon planning. Even recently emerged test-time scaling frameworks often resort to myopic, single-step lookaheads, which struggle to resolve the delayed feedback inherent in complex, occluded spatial environments. To address this limitation, we propose ParallelWorld, a multi-horizon test-time scaling framework for embodied reasoning. Instead of greedy, single-step trials, ParallelWorld empowers agents to simulate and evaluate multi-step future trajectories in parallel before committing to an action. Specifically, we introduce a verifier-guided tree-search paradigm. Starting from the current state, ParallelWorld branches into multiple parallel trajectories and rolls them out continuously across a multi-step horizon. At each simulation step, a verifier agent evaluates the intermediate state transitions, dynamically pruning unpromising branches and prioritizing paths with the highest information gain. Once the multi-step prospective simulation is complete, the agent synthesizes the long-horizon outcomes to commit to the optimal action sequence. Finally, an answer agent performs reasoning over the selected trajectory to produce the final reasoning. Extensive experiments on ESI-Bench demonstrate that ParallelWorld consistently improves active perception and reasoning performance.
Problem

Research questions and friction points this paper is trying to address.

Embodied Reasoning
Long-horizon Planning
Test-time Scaling
Delayed Feedback
Occluded Spatial Environments
Innovation

Methods, ideas, or system contributions that make the work stand out.

multi-horizon test-time scaling
verifier-guided tree-search
parallel trajectories
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
M
Min Chen
Tsinghua University
Shengjun Zhang
Shengjun Zhang
Tsinghua University
Y
Yuxin Li
Tsinghua University
Z
Zhang Zhang
Tsinghua University
Xin Fei
Xin Fei
National University of Singapore
Robotic ManipulationComputer Vision
C
Chong Xia
Tsinghua University
Y
Yueqi Duan
Tsinghua University