CSWAM: Better Causal Semantic Representations for Out-of-Distribution Generalization in World Action Models

📅 2026-09-16
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决FastWAM在视觉分布变化下的泛化问题,本文提出CSWAM模型,利用V-JEPA 2.1构建因果语义专家来增强表示能力,提高对未知场景和物体的适应性。
📝 Abstract
FastWAM-style world action models enable efficient action-only inference, but generalize poorly under visual distribution shifts. Their reconstruction-oriented representations emphasize appearance-specific details, limiting generalization to unseen scenes and objects. Without observation history, the model also lacks temporal evidence for robustly identifying task-relevant state changes and motion in unfamiliar visual conditions. To address these limitations, we present the Causal Semantic World Action Model (CSWAM), which augments FastWAM with a causal semantic expert built on V-JEPA 2.1. V-JEPA provides temporally grounded representations of semantic state changes and motion with less dependence on appearance-specific details. The expert learns their future evolution from a sparse history of current and past observations and shares the history-derived context with both the video and action streams through causal attention. At inference, CSWAM conditions action denoising on the current video state and observed semantic history, retaining efficient action-only inference. We conduct simulation and real-robot experiments to evaluate generalization under distribution shifts. With embodied pretraining, CSWAM raises Randomized success on RoboTwin 2.0 Clean-to-Randomized transfer from 10.16% to 45.18%, a gain of 35.02 percentage points over FastWAM. Across two real-robot tasks and three OOD difficulty levels, CSWAM improves average success over FastWAM by 42.5 percentage points, from 27.5% to 70.0%.
Problem

Research questions and friction points this paper is trying to address.

visual distribution shifts
generalization
temporal evidence
Innovation

Methods, ideas, or system contributions that make the work stand out.

Causal Semantic Representation
V-JEPA 2.1
Temporal Grounding
Causal Attention
🔎 Similar Papers