4D-WAM: Infusing Spatiotemporal Awareness into World Action Models through Trajectory Fields

📅 2026-08-08
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Existing world action models operate in 2D pixel space, struggling to effectively capture the 3D dynamic structures essential for robotic tasks, thereby creating a significant representational gap. To address this limitation, this work proposes 4D-WAM—a model-agnostic training strategy that introduces, for the first time, a trajectory-field-driven representation alignment mechanism. This approach jointly optimizes motion alignment—synchronizing feature changes across adjacent frames—and goal alignment—minimizing discrepancies in attention distributions between source and target frames—to achieve local 4D perception guided by long-horizon objectives. Experimental results demonstrate that 4D-WAM substantially enhances spatial understanding, execution accuracy, robustness, generalization, and overall versatility across multiple foundational models.
📝 Abstract
Building on recent advances in world models, World Action Models (WAMs) jointly model video prediction and action generation. However, they typically represent videos in 2D pixel space, creating a representation gap with 3D space in which robotic actions are executed. Recent 3D approaches introduce 3D information, but fail to fully exploit the dynamics of 3D structures. In this work, we propose 4D-WAM, a model-agnostic training strategy that injects spatiotemporal knowledge from 3D trajectory fields into WAMs through representation alignment. To this end, we introduce two complementary objectives: 1) motion alignment, which aligns temporal feature variations across adjacent frames and encourages the model to build local 4D awareness during training, and 2) destination alignment, which guides the model to infer the final destination from the source frame by minimizing the gap between their attention-like similarity distributions. Together, these objectives provide both local motion supervision and long-horizon goal guidance, enabling WAMs to learn trajectory-level spatiotemporal representations. Extensive in-distribution and out-of-distribution experiments across different base models demonstrate the model's improvements in spatial understanding, execution precision, robustness, generalization, and versatility.
Problem

Research questions and friction points this paper is trying to address.

World Action Models
spatiotemporal awareness
3D trajectory fields
representation gap
robotic actions
Innovation

Methods, ideas, or system contributions that make the work stand out.

4D-WAM
trajectory fields
spatiotemporal awareness
representation alignment
World Action Models