LM-X: Explainable Action Modeling with Progress, Event, and Uncertainty Prediction for Generalist Robot Manipulation

📅 2026-08-26
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出LM-X模型,通过预测任务进展、事件和不确定性来解决通用机器人操作中的长周期行为学习问题,提高控制的可解释性和成功率。
📝 Abstract
Generalist vision--language--action (VLA) policies learn long-horizon behavior mainly through short-horizon action prediction and reveal little beyond sampled commands. This creates two coupled bottlenecks: a single action target must implicitly absorb task progress, intermediate intent, and local reliability, while these control states remain hidden during execution. Inspired by functional principles of biological sensorimotor control, we introduce LM-X , which organizes prediction across task, event, and motor scales without claiming anatomical correspondence. Three explicitly supervised signals are emitted online and directly condition action generation: return-to-go (RTG) measures visible task progress, event-to-go (ETG) identifies the next semantic transition, and heteroscedastic action flow estimates local reliability through propagated variance. Explanation is therefore intrinsic to control rather than generated post hoc. Before a costly 20-day pretraining run on 64 NVIDIA B200 GPUs, a controlled five-task pretraining gate verifies the design: the complete model improves success by 16.0 points over the action-only backbone and by 10.8 points over the strongest single-head variant. We then train LM-X on more than 20,000 hours of real-robot trajectories, including over 1,000 hours of failed policy rollouts. LM-X achieves 74.1\% across 50 randomized-hard RoboTwin2.0 tasks versus 55.4\% for GR00T N1.7, and 68.6\% versus 50.7\% across seven real-robot tasks. RTG tracks semantic progress and visible regression, while variance rises during hesitation and oscillatory control. These results show that explicit multi-timescale predictive state can strengthen control while exposing interpretable internal estimates.
Problem

Research questions and friction points this paper is trying to address.

Generalist VLA policies
short-horizon action prediction
task progress
intermediate intent
local reliability
Innovation

Methods, ideas, or system contributions that make the work stand out.

LM-X
multi-timescale prediction
explainable control
return-to-go (RTG)
event-to-go (ETG)
💼 Related Jobs
No related jobs found.
J
Jin Lou
Humanoid Robot (Shanghai) Co., Ltd.
J
Jingxuan Zhu
E-surfing Digital Life Technology Co., Ltd., China Telecom
A
Andong Chen
Humanoid Robot (Shanghai) Co., Ltd.
X
Xupeng Wang
Humanoid Robot (Shanghai) Co., Ltd.
Y
Yuan Xu
Humanoid Robot (Shanghai) Co., Ltd.
Y
Yuexuan Li
Humanoid Robot (Shanghai) Co., Ltd.
X
Xingdong Zhu
Humanoid Robot (Shanghai) Co., Ltd.
Z
Zhijie Zhu
Humanoid Robot (Shanghai) Co., Ltd.
Y
Yingwei Ji
Humanoid Robot (Shanghai) Co., Ltd.
W
Wenpeng Nie
Humanoid Robot (Shanghai) Co., Ltd.
J
Jingyi Li
E-surfing Digital Life Technology Co., Ltd., China Telecom
Liangliang Chen
Liangliang Chen
Georgia Institute of Technology
Machine LearningRoboticsHuman-in-the-loop ControlAI in EducationControl Theory & Application
J
Jinyan Liu
E-surfing Digital Life Technology Co., Ltd., China Telecom
Z
Zhiqi Song
E-surfing Digital Life Technology Co., Ltd., China Telecom
J
Jidong Zhang
E-surfing Digital Life Technology Co., Ltd., China Telecom
H
Hongming Li
E-surfing Digital Life Technology Co., Ltd., China Telecom
Yuchen Zhu
Yuchen Zhu
Georgia Institute of Technology
Diffusion ModelsDiscrete DiffusionVision-language Model