PV-WM: A Heterogeneous Micro-Macro World Model for Articulated Pedestrian-Vehicle Co-Rollout

📅 2026-09-07
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出PV-WM模型,通过同步异构状态递归推进行人和车辆的预测,解决了行人-车辆联合预测中物理尺度差异的问题。
📝 Abstract
Local pedestrian-vehicle forecasting spans heterogeneous physical scales: pedestrians combine root locomotion with articulated motion, whereas vehicles are rigid bodies described by kinematic state and oriented extent. Existing road-agent forecasters typically omit pedestrian articulation, while pose forecasters leave vehicle futures outside the learned rollout. We introduce PV-WM, a history-only world model over structured post-perception tracks. It recurrently advances pedestrian root motion, 15-joint articulation, and learned vehicle states within a synchronized heterogeneous state. The generated pedestrian and vehicle chunks supply the next recurrent boundary; vehicle boxes are reconstructed from predicted center and heading with observed extent, and P-V geometry is recomputed after every transition. Relative to a matched one-shot complete-state predictor, recurrent execution reduces Root ADE by 12.7% and MPJPE by 14.8%. Feedback interventions show that later predictions depend on the content, temporal order, and pedestrian identity of generated articulation. Across 824 aligned Waymo contexts, with 797 providing valid future vehicle support, PV-WM reduces Root ADE by 5.2%, MPJPE by 7.6%, P-V distance error by 11.9%, and oriented-box closest-approach error by 5.8% relative to a validation-selected Modular Specialist. The single-network model uses 57.1% fewer parameters, 96.5% lower average FLOPs per local scene, and 25.5% lower measured p95 latency. PV-WM unifies this heterogeneous future state while preserving type-specific pedestrian and vehicle dynamics.
Problem

Research questions and friction points this paper is trying to address.

pedestrian-vehicle forecasting
articulated motion
heterogeneous physical scales
Innovation

Methods, ideas, or system contributions that make the work stand out.

heterogeneous state
recurrent execution
articulated motion
vehicle state
synchronized
H
Haozhuang Chi
Nanyang Technological University, Singapore
Jingsong Liang
Jingsong Liang
National University of Singapore
Robot LearningReinforcement LearningMulti-Agent SystemsEmbodied Intelligence
Ziying Song
Ziying Song
Beijing Jiaotong University
Object DetectionComputer VisionDeep Learning
L
Lei Yang
Nanyang Technological University, Singapore
S
Shihao Li
Beijing Institute of Technology, Beijing, China
H
Haoruo Zhang
Nanyang Technological University, Singapore
C
Chen Lv
Nanyang Technological University, Singapore