Learning to Use Imagination: Progress-Conditioned Future Utilization for World Action Models

📅 2026-09-06
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文针对WAM中未来想象适应性不足的问题,提出ProWAM模型,通过引入执行进度作为中介表示来优化想象利用。
📝 Abstract
World Action Models (WAMs) extend Vision-Language-Action (VLA) models by incorporating future visual dynamics into action generation. However, existing WAMs often utilize imagined futures with limited adaptation to evolving execution progress, potentially introducing distracting or unreliable predictive cues. This limitation arises from two empirically identified forms of non-uniformity in future utility: (i) at the inter-progress level, the utility of imagined futures varies across execution stages as control demands change; and (ii) at the intra-progress level, individual future latents exhibit heterogeneous relevance within the same progress state. To address these limitations, we propose ProWAM, a Progress-Conditioned World Action Model that introduces execution progress as an explicit intermediate representation for adaptive imagination utilization. ProWAM comprises two tightly coupled components: (1) To obtain a reliable representation of execution progress, we propose the Self-Supervised Dual-Temporal Progress Encoder (SS-DTPE). SS-DTPE couples short-term action-observation interaction modeling with long-term recurrent progress aggregation to capture recent execution feedback and accumulated task history. (2) Conditioned on the progress representation from SS-DTPE, we propose the Hierarchical Progress-Conditioned Imagination Modulation (HPIM) to adapt imagination utilization to execution progress. HPIM operates at two complementary levels: an inter-progress global modulation mechanism adapts future utilization across execution stages, while an intra-progress relevance mechanism differentiates individual future latents within each progress state. Extensive experiments demonstrate consistent gains over strong VLA and WAM baselines.
Problem

Research questions and friction points this paper is trying to address.

World Action Models
execution progress
future utility
visual dynamics
imagination utilization
Innovation

Methods, ideas, or system contributions that make the work stand out.

Progress-Conditioned World Action Model
Self-Supervised Dual-Temporal Progress Encoder
Hierarchical Progress-Conditioned Imagination Modulation
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Y
Yijie Zhu
Harbin Institute of Technology (Shenzhen), 518055, China, and Great Bay University, Dongguan 523000, China
Zitong Yu
Zitong Yu
U.S. Food and Drug Administration
Medical imagingDeep learningMachine learningImage reconstruction
W
Wei Li
Harbin Institute of Technology (Shenzhen), 518055, China
H
Hui Ma
Great Bay University, Dongguan 523000, China
W
Wen Li
University of Electronic Science and Technology of China, Chengdu 611731, China
Rui Shao
Rui Shao
Professor, Harbin Institute of Technology (Shenzhen)
Computer VisionMultimodal LLMEmbodied AI
L
Liqiang Nie
Harbin Institute of Technology (Shenzhen), 518055, China