GloVLA: Let Geometry Move and Local VLA Interact for Robust Object-Centric Manipulation in Unstructured Environments

📅 2026-09-05
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出GloVLA框架,通过几何运输控制器和局部VLA策略解决非结构化环境中基于物体操作的鲁棒性问题。
📝 Abstract
Vision-language-action (VLA) models have shown promising generalization for language-conditioned robot manipulation, but deploying them in unstructured environments remains challenging. A single end-to-end VLA policy must simultaneously solve long-range transport of the end effector to task-relevant regions and short-horizon, contact-rich interaction upon arrival. This formulation is inefficient and brittle: small visual shifts, distractors, clutter, occlusions, or unfavorable initial gripper poses can push the policy outside the local state distribution in which it was trained, leading to task failure. We introduce GloVLA, a hybrid framework that explicitly separates object-centric manipulation into two complementary regimes: a geometric transport controller moves the end-effector into interaction-centric handoff regions, and local VLA policies handle only the short-horizon interaction phases. GloVLA is model-agnostic and can be integrated with different VLA backbones with no additional demonstrations and no changes to the action space or success predicate. Experiments on standard LIBERO and LIBERO-Plus Object tasks together with a newly introduced LIBERO-Challenge benchmark ettings with clutter, distractors,illumination changes, visual shifts, and obstruction show that GloVLA improves task success and substantially lowers VLA inference cost compared with full end-to-Challenge, full-trajectory GR00T N1.6execution degrades to 20.9% average success while GloVLA retains 88.5%; on a physical UR10e, overall success improves from 35.6% to 90.0% while mean inference time is more than halved. Videos and additional results are available at https://glovla-project.github.io/
Problem

Research questions and friction points this paper is trying to address.

Vision-language-action (VLA) models
unstructured environments
object-centric manipulation
long-range transport
short-horizon interaction
Innovation

Methods, ideas, or system contributions that make the work stand out.

geometric transport controller
local VLA policies
object-centric manipulation
unstructured environments
hybrid framework
💼 Related Jobs
No related jobs found.
T
Truong Thanh Nguyen
VinRobotics, Vietnam
H
Huy Hoang Nguyen
Austrian Institute of Technology, Vienna, Austria
H
Ha Anh Nguyen
Hanoi University of Science and Technology, Vietnam
B
Binh Khanh Dinh
VinRobotics, Vietnam
Ngo Anh Vien
Ngo Anh Vien
VinRobotics & VinUni, ex-BCAI
machine learningrobotics
D
Duy Nguyen Ho Minh
German Research Center for Artificial Intelligence (DFKI), Germany; University of Stuttgart, Germany; International Max Planck Research School for Intelligent Systems (IMPRS-IS), Germany
Minh Nhat Vu
Minh Nhat Vu
Automation & Control Institute (ACIN), Vienna, Austria
Robotics
Ngan Le
Ngan Le
University of Arkansas
Artificial IntelligenceMachine LearningComputer Vision