VLBiMan++: Expanding the Generalization Boundary of Vision-Language Anchored One-Shot Bimanual Manipulation

📅 2026-09-13
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出VLBiMan++框架,通过视觉-语言锚定的一次性双臂操作方法,解决多样化任务、对象、场景、机器人平台及执行条件下的泛化问题。
📝 Abstract
Generalizable bimanual robotic manipulation requires a reusable task prior that can persist across increasingly diverse tasks, objects, scenes, embodiments, and execution conditions, thus avoiding the prohibitive cost of large-scale teleoperated demonstrations and policy retraining. In this work, we present VLBiMan++, an extended framework that expands the generalization boundary of vision-language anchored one-shot bimanual manipulation. Starting from a single human demonstration, VLBiMan++ performs task-aware decomposition to identify reusable and adaptable skill components, and employs vision-language grounded geometric adaptation to transfer these skills to novel configurations without retraining. Building on this foundation, we systematically extend generalization along five dimensions: task generalization through diverse and long-horizon skill compositions; object generalization across unseen categories, varying geometries, and more complex articulated or deformable objects; scene generalization under clutter, occlusion, and dynamic interference; embodiment generalization across heterogeneous dual-arm robotic platforms; and deployment generalization through prolonged closed-loop execution under repeated external perturbations. To support this broader scope, we further introduce object-state-aware adaptation and lightweight trajectory optimization mechanisms that accommodate changes beyond simple rigid 6-DoF pose variations while preserving reliable bimanual coordination. Extensive real-world experiments demonstrate that VLBiMan++ maintains strong task success and adaptation capability across these increasingly challenging settings. Overall, VLBiMan++ advances one-shot bimanual manipulation from demonstrating isolated transferability toward a more systematic and scalable framework for generalization across tasks, objects, scenes, embodiments, and long-term deployment conditions.
Problem

Research questions and friction points this paper is trying to address.

bimanual robotic manipulation
generalization
vision-language anchored
one-shot learning
task prior
Innovation

Methods, ideas, or system contributions that make the work stand out.

task-aware decomposition
vision-language grounded geometric adaptation
object-state-aware adaptation
lightweight trajectory optimization
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
H
Huayi Zhou
Guangdong Provincial Key Laboratory of Visual Media and Multidimensional Intelligence, College of Computer Science and Software Engineering, Shenzhen University
W
Wei Gao
DexForce, Shenzhen
Y
Yiyang Han
Faculty of Engineering, Department of Computing, Imperial College London
K
Kui Jia
School of Data Science, The Chinese University of Hong Kong, Shenzhen; DexForce, Shenzhen
Hui Huang
Hui Huang
Chair Professor and CS Dean, Shenzhen University
GraphicsGeometryPointsShapesImages