What Matters for Latent Actions in Robot Learning

📅 2026-08-19
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过统一框架评估了41种潜在动作模型设计选择,解决了机器人学习中潜在动作效果不一致的问题,并发现微调视觉-语言模型能有效提升下游策略学习。
📝 Abstract
Latent Action Models (LAMs) have emerged as a promising paradigm for enabling robot learning to leverage large-scale unlabeled videos through latent actions that serve as compact surrogates for physical actions. Despite rapid progress, research on LAM remains highly fragmented, with existing methods evaluating different design choices in isolation under inconsistent experimental settings, making it difficult to identify the factors that truly determine downstream robotic manipulation performance. In this work, we present the first comprehensive empirical study of latent action learning for robotic manipulation. We unify representative LAM methods within a common autoencoding framework and systematically investigate 41 LAM design choices across three dimensions, including latent action modeling paradigms, learning objectives and regularization methods, and latent action integration strategies. We further examine four proxy metrics for evaluating latent action quality and assess their ability to reliably predict downstream robotic manipulation performance. Extensive experiments on three widely used benchmarks provide strong empirical evidence that fine-tuning vision-language model (VLM) backbones with latent actions provides a stronger initialization for downstream policy learning, with further validation on real-world robot manipulation tasks.
Problem

Research questions and friction points this paper is trying to address.

Latent Action Models
robot learning
unlabeled videos
manipulation performance
design choices
Innovation

Methods, ideas, or system contributions that make the work stand out.

Latent Action Models
robotic manipulation
autoencoding framework
vision-language model backbones
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
X
Xizhou Bu
College of Intelligent Robotics and Advanced Manufacturing, Fudan University, Shanghai, 200433, China
Q
Qingda Hu
College of Intelligent Robotics and Advanced Manufacturing, Fudan University, Shanghai, 200433, China
L
Lei Zhou
Xiaomi EV, Beijing, 100085, China
Lingfeng Zhang
Lingfeng Zhang
PhD student at Tsinghua University
embodied ai
Yingbo Tang
Yingbo Tang
Institute of Automation,Chinese Academy of Sciences
Z
Zihao Liu
Guangdong Provincial Key Laboratory of Computility Microelectronics, Faculty of Computility Microelectronics, Shenzhen University of Advanced Technology, Shenzhen, 518107, China
X
Xinyi Tao
School of Aeronautics and Astronautics, Sichuan University, Chengdu, 610207, China
Zhiqiang Ma
Zhiqiang Ma
JPMorgan Chase
topic modelingtext miningmachine learning
Qingqiu Huang
Qingqiu Huang
Research Engineer,Huawei
computer vision,autonomous driving
C
Chufeng Tang
Morphi Intelligence Technology Co., Ltd., Shenzhen, 518063, China
H
Hongbo Wang
College of Intelligent Robotics and Advanced Manufacturing, Fudan University, Shanghai, 200433, China
J
Jing Zhang
School of Artificial Intelligence, Wuhan University, Wuhan, 430072, China
Jiayi Ma
Jiayi Ma
Wuhan University
Computer VisionImage FusionImage Matching
H
Hangjun Ye
Xiaomi EV, Beijing, 100085, China
Wei Li
Wei Li
Fudan University
RoboticsCollective IntelligenceMachine Learning
Xiaoshuai Hao
Xiaoshuai Hao
Beijing Academy of Artificial Intelligence,BAAI
vision and language