TemporalFlow-VLA: Learning Physically Grounded Execution History for Long-Horizon Robot Manipulation

📅 2026-08-27
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决多阶段操作中物理变化捕捉问题,提出TemporalFlow-VLA方法,通过学习紧凑的执行历史和基于物理的时间监督来提高机器人长期操作的成功率。
📝 Abstract
Vision-language-action (VLA) models leverage pretrained vision-language representations for robot control, yet simply adding historical frames does not reliably capture recent physical change. This is especially problematic in multi-stage manipulation, where visually similar states may require different actions depending on prior execution. To address this challenge, we present TemporalFlow-VLA, which learns compact execution history through physically grounded temporal supervision. Using recorded robot states, robot geometry, and calibrated cameras, we construct robot-surface temporal flow as a training-only target and supervise two execution-aligned temporal queries that provide structured history to the action expert. The geometric supervision path is not evaluated at deployment. TemporalFlow-VLA achieves 97.63 +/- 0.26% average success on LIBERO, including 96.60 +/- 0.87% on LIBERO Long, and 85.5%/84.2% Clean/Randomized success across 12 RoboTwin tasks. It shows its clearest advantage over prior methods on longer-horizon, multi-stage manipulation. Controlled history interventions show that action prediction depends on both historical content and temporal order. With asynchronous feature caching, temporal conditioning maintains single-frame-level server-side sampling latency without additional historical-encoding overhead. Overall, TemporalFlow-VLA provides a compact, physically grounded interface for exploiting ordered execution history without explicit motion estimation or geometric processing at deployment.
Problem

Research questions and friction points this paper is trying to address.

vision-language-action
robot control
historical frames
multi-stage manipulation
physical change
Innovation

Methods, ideas, or system contributions that make the work stand out.

TemporalFlow-VLA
physically grounded temporal supervision
execution history
multi-stage manipulation
asynchronous feature caching
💼 Related Jobs
No related jobs found.
J
Jiarui Yang
The Hong Kong University of Science and Technology (Guangzhou), Guangzhou, China.
Yehao Lu
Yehao Lu
Zhejiang University
Autonomous Driving3D ReconstructionSwarm Robot
Y
Yuning Su
Simon Fraser University, Burnaby, BC, Canada.
Y
Yu Zhong
AgiBot, Shanghai, China.
Y
Yufeng Xie
AgiBot, Shanghai, China.
Yazhou Zhang
Yazhou Zhang
Associate Professor, Tianjin University
Sentiment AnalysisQuantum CognitionSarcasm DetectionHumor Analysis
Haiyu Lan
Haiyu Lan
University of Calgary
PerceptionLocalization and Mapping for Autonomous VehicleRobotics.
K
Kaixiang Lu
AgiBot, Shanghai, China.
P
Peiwen Lin
AgiBot, Shanghai, China.
C
Chuang Wang
AgiBot, Shanghai, China.
Junwei Liang
Junwei Liang
Assistant Professor, HKUST (Guangzhou) | CSE, HKUST | Ph.D. @CMU
Computer VisionRoboticsEmbodied AITrajectory Prediction
E
Enyu Li
AgiBot, Shanghai, China.