VLBiMan++: Expanding the Generalization Boundary of Vision-Language Anchored One-Shot Bimanual Manipulation
本文提出VLBiMan++框架,通过视觉-语言锚定的一次性双臂操作方法,解决多样化任务、对象、场景、机器人平台及执行条件下的泛化问题。
本文提出VLBiMan++框架,通过视觉-语言锚定的一次性双臂操作方法,解决多样化任务、对象、场景、机器人平台及执行条件下的泛化问题。
This work addresses the limited generalization of real-world robotic manipulation policies, which stems from the scarcity and insufficient diversity of real data. To systematically advance research on generalizable manipulation across tasks and environments, the paper introduces the first unified benchmark that integrates large-scale synthetic data training with standardized real-robot evaluation. Methodologically, it pioneers the combination of synthetic skill learning with rigorous real-world validation, leveraging state-of-the-art architectures—including Transformers, diffusion models, and vision-language-action frameworks—to construct manipulation policies. The project provides multiple baseline implementations and establishes a reproducible, comparable evaluation framework for generalization, thereby facilitating the development of data-efficient and adaptable general-purpose robotic manipulation systems.
This work addresses the limitations of existing speech-driven motion generation methods for humanoid robots, which rely on human-centric motion representations and often produce physically infeasible or under-expressive actions due to embodiment mismatches. To ensure embodiment consistency, the authors propose an embodiment-aware, native motion generation framework that directly predicts robot joint trajectories from speech in an end-to-end manner, bypassing intermediate human motion representations. The core innovations include an IK-EER motion filtering mechanism and the PhysDrift generative model, which jointly integrate inverse kinematics optimization, prosody-aligned speech conditioning, and physics-based regularization to optimize both speech-motion synchronization and physical plausibility. Experiments demonstrate significant improvements in motion alignment, naturalness, smoothness, and inference efficiency, with successful real-time deployment on a physical humanoid robot.
Existing end-to-end robotic manipulation approaches rely on 2D visual inputs, which struggle to capture the inherently 3D nature of tasks and suffer from misalignment between perception and action spaces in both spatial and temporal dimensions, limiting generalization. This work proposes a pixel-wise 3D visual representation that constructs aligned vertex maps using camera calibration and depth information, unifying multi-view perception and robot actions within a shared world coordinate frame. To achieve viewpoint-invariant encoding, a bird’s-eye-view (BEV) representation is introduced, complemented by a cross-platform trajectory time-alignment mechanism. The proposed approach substantially mitigates spatiotemporal misalignment between perception and action, significantly enhancing policy generalization and robustness across diverse robots, viewpoints, and human demonstrators. The authors also release pretrained models, code, and a complete data processing pipeline to support reproducibility and further research.
This work addresses the vulnerability of robotic manipulation tasks to failures in dynamic, unstructured environments, where existing approaches rely on post-hoc detection and reactive recovery, resulting in high latency and insufficient robustness. To overcome these limitations, the authors propose AgentChord, a novel multi-agent system that incorporates human-inspired prospective planning into robotic fault recovery for the first time. AgentChord structures the recovery process around three specialized agents—Composer, Orchestrator, and Conductor—and models tasks as directed task graphs with pre-embedded, context-aware recovery branches. By integrating low-latency monitoring with precompiled recovery strategies, the system enables proactive anticipation and immediate response without requiring online replanning. Experimental results demonstrate that AgentChord significantly improves success rates and execution efficiency across diverse long-horizon bimanual manipulation tasks, thereby enhancing robotic reliability and autonomy in real-world settings.
本文提出VLBiMan++框架,通过视觉-语言锚定的一次性双臂操作方法,解决多样化任务、对象、场景、机器人平台及执行条件下的泛化问题。
This work addresses the limited generalization of real-world robotic manipulation policies, which stems from the scarcity and insufficient diversity of real data. To systematically advance research on generalizable manipulation across tasks and environments, the paper introduces the first unified benchmark that integrates large-scale synthetic data training with standardized real-robot evaluation. Methodologically, it pioneers the combination of synthetic skill learning with rigorous real-world validation, leveraging state-of-the-art architectures—including Transformers, diffusion models, and vision-language-action frameworks—to construct manipulation policies. The project provides multiple baseline implementations and establishes a reproducible, comparable evaluation framework for generalization, thereby facilitating the development of data-efficient and adaptable general-purpose robotic manipulation systems.
This work addresses the limitations of existing speech-driven motion generation methods for humanoid robots, which rely on human-centric motion representations and often produce physically infeasible or under-expressive actions due to embodiment mismatches. To ensure embodiment consistency, the authors propose an embodiment-aware, native motion generation framework that directly predicts robot joint trajectories from speech in an end-to-end manner, bypassing intermediate human motion representations. The core innovations include an IK-EER motion filtering mechanism and the PhysDrift generative model, which jointly integrate inverse kinematics optimization, prosody-aligned speech conditioning, and physics-based regularization to optimize both speech-motion synchronization and physical plausibility. Experiments demonstrate significant improvements in motion alignment, naturalness, smoothness, and inference efficiency, with successful real-time deployment on a physical humanoid robot.
Existing end-to-end robotic manipulation approaches rely on 2D visual inputs, which struggle to capture the inherently 3D nature of tasks and suffer from misalignment between perception and action spaces in both spatial and temporal dimensions, limiting generalization. This work proposes a pixel-wise 3D visual representation that constructs aligned vertex maps using camera calibration and depth information, unifying multi-view perception and robot actions within a shared world coordinate frame. To achieve viewpoint-invariant encoding, a bird’s-eye-view (BEV) representation is introduced, complemented by a cross-platform trajectory time-alignment mechanism. The proposed approach substantially mitigates spatiotemporal misalignment between perception and action, significantly enhancing policy generalization and robustness across diverse robots, viewpoints, and human demonstrators. The authors also release pretrained models, code, and a complete data processing pipeline to support reproducibility and further research.
This work addresses the vulnerability of robotic manipulation tasks to failures in dynamic, unstructured environments, where existing approaches rely on post-hoc detection and reactive recovery, resulting in high latency and insufficient robustness. To overcome these limitations, the authors propose AgentChord, a novel multi-agent system that incorporates human-inspired prospective planning into robotic fault recovery for the first time. AgentChord structures the recovery process around three specialized agents—Composer, Orchestrator, and Conductor—and models tasks as directed task graphs with pre-embedded, context-aware recovery branches. By integrating low-latency monitoring with precompiled recovery strategies, the system enables proactive anticipation and immediate response without requiring online replanning. Experimental results demonstrate that AgentChord significantly improves success rates and execution efficiency across diverse long-horizon bimanual manipulation tasks, thereby enhancing robotic reliability and autonomy in real-world settings.