Gripper-aware Vision Language Action Models
为解决现有视觉语言动作模型忽视夹爪类型差异的问题,通过构建多夹爪数据集MiGA和提出GVLA方法来提升机器人对不同夹爪的适应性和操作性能。
为解决现有视觉语言动作模型忽视夹爪类型差异的问题,通过构建多夹爪数据集MiGA和提出GVLA方法来提升机器人对不同夹爪的适应性和操作性能。
Existing reinforcement learning methods often suffer from training instability and poor scalability in high-dimensional action spaces. This work proposes the QGF algorithm, which uniquely performs policy optimization entirely at test time: it first pretrains a highly expressive flow-based policy via behavioral cloning and then, during inference, leverages value gradients from a critic to guide the flow model toward generating higher-value actions—eliminating the need for additional policy training. Evaluated across multiple offline reinforcement learning benchmarks, QGF significantly outperforms existing test-time optimization approaches, matching or surpassing state-of-the-art training-time algorithms while incurring lower computational overhead, thereby achieving a favorable balance between stability and efficiency.
This work addresses the challenge of transferring vision–language–action (VLA) models, pretrained on fixed-base robotic platforms, to highly dynamic and underactuated aerial manipulation systems, where dynamics mismatch severely degrades performance. To bridge this gap, the authors propose a payload-aware guidance mechanism that injects physical constraints during inference, alongside a synthetic navigation dataset generated via Gaussian splatting to alleviate real-world data scarcity. Notably, the approach enables effective zero-shot transfer to aerial grasping and navigation tasks without fine-tuning the base VLA model. Extensive real-world evaluation across 460 trials demonstrates that the synthetic data boosts navigation success from 81% to 100%, while payload-aware guidance increases grasping success from 23% to 50%. The integrated system achieves a 62% success rate on long-horizon compositional tasks.
This work addresses the challenge of jointly modeling long-term semantic memory and short-term perceptual memory in end-to-end robot learning for complex, multi-stage tasks. The authors propose a multi-scale embodied memory architecture that, for the first time, integrates multimodal and multi-granularity memory mechanisms into robotic policy learning. Specifically, a video encoder compresses short-term visual memory to handle occlusions, while a language model processes long-term semantic memory represented in textual form. These components are unified within a vision–language–action policy framework. The approach successfully executes long-horizon tasks—such as kitchen cleaning and sandwich preparation—lasting up to fifteen minutes, demonstrating the ability to adapt manipulation strategies based on contextual cues.
This work aims to achieve general-purpose robotic policies capable of zero-shot deployment across diverse morphologies without morphology-specific fine-tuning. To this end, the authors propose Language-Action Pretraining (LAP), a method that represents low-level actions as natural language tokens, aligning their input-output distribution with that of vision-language models. By unifying action prediction and visual question answering within a co-training framework, LAP enables significant zero-shot transfer to unseen robot morphologies—without requiring custom tokenizers, costly annotations, or morphology-specific architectures. The resulting LAP-3B model achieves an average zero-shot success rate exceeding 50% across multiple novel robots and manipulation tasks, representing approximately a two-fold improvement over the current state-of-the-art vision-language-action (VLA) models.
为解决现有视觉语言动作模型忽视夹爪类型差异的问题,通过构建多夹爪数据集MiGA和提出GVLA方法来提升机器人对不同夹爪的适应性和操作性能。
Existing reinforcement learning methods often suffer from training instability and poor scalability in high-dimensional action spaces. This work proposes the QGF algorithm, which uniquely performs policy optimization entirely at test time: it first pretrains a highly expressive flow-based policy via behavioral cloning and then, during inference, leverages value gradients from a critic to guide the flow model toward generating higher-value actions—eliminating the need for additional policy training. Evaluated across multiple offline reinforcement learning benchmarks, QGF significantly outperforms existing test-time optimization approaches, matching or surpassing state-of-the-art training-time algorithms while incurring lower computational overhead, thereby achieving a favorable balance between stability and efficiency.
This work addresses the challenge of transferring vision–language–action (VLA) models, pretrained on fixed-base robotic platforms, to highly dynamic and underactuated aerial manipulation systems, where dynamics mismatch severely degrades performance. To bridge this gap, the authors propose a payload-aware guidance mechanism that injects physical constraints during inference, alongside a synthetic navigation dataset generated via Gaussian splatting to alleviate real-world data scarcity. Notably, the approach enables effective zero-shot transfer to aerial grasping and navigation tasks without fine-tuning the base VLA model. Extensive real-world evaluation across 460 trials demonstrates that the synthetic data boosts navigation success from 81% to 100%, while payload-aware guidance increases grasping success from 23% to 50%. The integrated system achieves a 62% success rate on long-horizon compositional tasks.
This work addresses the challenge of jointly modeling long-term semantic memory and short-term perceptual memory in end-to-end robot learning for complex, multi-stage tasks. The authors propose a multi-scale embodied memory architecture that, for the first time, integrates multimodal and multi-granularity memory mechanisms into robotic policy learning. Specifically, a video encoder compresses short-term visual memory to handle occlusions, while a language model processes long-term semantic memory represented in textual form. These components are unified within a vision–language–action policy framework. The approach successfully executes long-horizon tasks—such as kitchen cleaning and sandwich preparation—lasting up to fifteen minutes, demonstrating the ability to adapt manipulation strategies based on contextual cues.
This work aims to achieve general-purpose robotic policies capable of zero-shot deployment across diverse morphologies without morphology-specific fine-tuning. To this end, the authors propose Language-Action Pretraining (LAP), a method that represents low-level actions as natural language tokens, aligning their input-output distribution with that of vision-language models. By unifying action prediction and visual question answering within a co-training framework, LAP enables significant zero-shot transfer to unseen robot morphologies—without requiring custom tokenizers, costly annotations, or morphology-specific architectures. The resulting LAP-3B model achieves an average zero-shot success rate exceeding 50% across multiple novel robots and manipulation tasks, representing approximately a two-fold improvement over the current state-of-the-art vision-language-action (VLA) models.