Institution profile

Astribot

Industry researchasia · cn
Official website
Research library18linked papers
Opportunities0open roles
Selected work

Representative Papers

CLAP: Contrastive Latent Action Pretraining for Learning Vision-Language-Action Models from Human Videos

Jan 07, 2026arXiv.org

This work addresses the limitations of current vision-language-action (VLA) models, which are hindered by the scarcity of robotic demonstration data and the susceptibility of human-video-based latent action approaches to visual distractions, impeding the extraction of executable skills. To overcome these challenges, we propose the Contrastive Latent Action Pretraining (CLAP) framework, which aligns the visual latent space of human videos with robot proprioceptive trajectories through contrastive learning and maps actions to an executable quantized codebook. We introduce a dual-branch VLA architecture—comprising CLAP-NTP and CLAP-RF—that integrates Rectified Flow with knowledge-matching regularization to effectively mitigate catastrophic forgetting during fine-tuning. Experiments demonstrate that our approach significantly outperforms baseline methods in skill transfer tasks, achieving superior instruction following, object generalization, and high-frequency precise manipulation capabilities.

1 citationsRead paper

Evolve Vision-Language-Action Model into an Agent with On-the-fly Tool-use

Aug 14, 2026

This study addresses the limited generalization and high data dependency of Vision-Language-Action (VLA) models by proposing ART, a framework that integrates tool use into VLAs to reduce action space complexity. By combining few-shot dataset construction with long-horizon reasoning training, ART achieves modular model enhancement. Experiments demonstrate that ART outperforms mainstream baselines by 20% in success rate across both simulated and real-world tasks while significantly reducing data requirements. The approach markedly improves robotic generalization in complex environments and enables lightweight, efficient deployment. Ultimately, this work establishes a scalable paradigm for embodied intelligence, effectively mitigating key bottlenecks in current VLA architectures through strategic tool integration and optimized training methodologies.

0 citationsRead paper

PhysMani: Physics-principled 3D World Model for Dynamic Object Manipulation

Jul 02, 2026

Manipulating fast-moving dynamic objects in unstructured 3D environments remains highly challenging, as existing vision-language-action and world model approaches struggle to accurately capture 3D geometry and generate physically plausible future states. This work proposes a novel framework that integrates physical priors into a 3D Gaussian world model coupled with a forward-looking policy architecture. It introduces, for the first time, a divergence-free Gaussian velocity field, optimized online to ensure physically consistent dynamics prediction. Furthermore, a cross-attention mechanism with learnable tokens seamlessly incorporates these predicted dynamics into a unified vision-language-action policy. The authors also introduce PhysMani-Bench, a new benchmark comprising 16 diverse tasks. Experiments demonstrate substantial improvements over strong baselines, with significantly higher task success rates in both simulation and real-world robotic settings.

0 citationsRead paper

StableVLA: Towards Robust Vision-Language-Action Models without Extra Data

May 18, 2026

This work addresses the limited robustness of vision-language-action models under unseen real-world visual perturbations by introducing the Information Bottleneck Adapter (IB-Adapter)—a lightweight, information-theoretic module that selectively filters potential noise from visual inputs. Without requiring additional data or augmentation strategies, IB-Adapter enhances model robustness while adding fewer than 10 million parameters. Remarkably, when integrated into small-scale models, it achieves performance comparable to that of 7B-parameter counterparts and yields an average 30% improvement under both synthetic and real-world disturbances. The approach significantly outperforms OpenPi while preserving accuracy on long-horizon tasks.

0 citationsRead paper

When to Trust Imagination: Adaptive Action Execution for World Action Models

May 07, 2026

This work addresses the limitations of existing World Action Models (WAMs), which rely on fixed-length action sequences and struggle to assess consistency between predicted futures and real-world observations, leading to poor responsiveness in contact-rich or complex tasks. The authors formulate adaptive execution as a future-reality consistency verification problem and introduce the Future Forward Dynamics Causal Attention (FFDC) mechanism. FFDC employs a lightweight verifier to dynamically adjust action chunk lengths and incorporates Mixture-of-Horizon Training to enhance coverage of long-horizon trajectories. The approach integrates predicted actions, visual dynamics, real observations, and language instructions, leveraging causal attention for consistency evaluation. Experiments demonstrate that the method reduces forward computation by 69.10%, shortens execution time by 34.02%, and improves success rates by 2.54% on RoboTwin, with a 35% absolute gain in real-robot trials.

0 citationsRead paper
Recent publications

Latest Papers

Evolve Vision-Language-Action Model into an Agent with On-the-fly Tool-use

Aug 14, 2026

This study addresses the limited generalization and high data dependency of Vision-Language-Action (VLA) models by proposing ART, a framework that integrates tool use into VLAs to reduce action space complexity. By combining few-shot dataset construction with long-horizon reasoning training, ART achieves modular model enhancement. Experiments demonstrate that ART outperforms mainstream baselines by 20% in success rate across both simulated and real-world tasks while significantly reducing data requirements. The approach markedly improves robotic generalization in complex environments and enables lightweight, efficient deployment. Ultimately, this work establishes a scalable paradigm for embodied intelligence, effectively mitigating key bottlenecks in current VLA architectures through strategic tool integration and optimized training methodologies.

0 citationsRead paper

PhysMani: Physics-principled 3D World Model for Dynamic Object Manipulation

Jul 02, 2026

Manipulating fast-moving dynamic objects in unstructured 3D environments remains highly challenging, as existing vision-language-action and world model approaches struggle to accurately capture 3D geometry and generate physically plausible future states. This work proposes a novel framework that integrates physical priors into a 3D Gaussian world model coupled with a forward-looking policy architecture. It introduces, for the first time, a divergence-free Gaussian velocity field, optimized online to ensure physically consistent dynamics prediction. Furthermore, a cross-attention mechanism with learnable tokens seamlessly incorporates these predicted dynamics into a unified vision-language-action policy. The authors also introduce PhysMani-Bench, a new benchmark comprising 16 diverse tasks. Experiments demonstrate substantial improvements over strong baselines, with significantly higher task success rates in both simulation and real-world robotic settings.

0 citationsRead paper

StableVLA: Towards Robust Vision-Language-Action Models without Extra Data

May 18, 2026

This work addresses the limited robustness of vision-language-action models under unseen real-world visual perturbations by introducing the Information Bottleneck Adapter (IB-Adapter)—a lightweight, information-theoretic module that selectively filters potential noise from visual inputs. Without requiring additional data or augmentation strategies, IB-Adapter enhances model robustness while adding fewer than 10 million parameters. Remarkably, when integrated into small-scale models, it achieves performance comparable to that of 7B-parameter counterparts and yields an average 30% improvement under both synthetic and real-world disturbances. The approach significantly outperforms OpenPi while preserving accuracy on long-horizon tasks.

0 citationsRead paper

When to Trust Imagination: Adaptive Action Execution for World Action Models

May 07, 2026

This work addresses the limitations of existing World Action Models (WAMs), which rely on fixed-length action sequences and struggle to assess consistency between predicted futures and real-world observations, leading to poor responsiveness in contact-rich or complex tasks. The authors formulate adaptive execution as a future-reality consistency verification problem and introduce the Future Forward Dynamics Causal Attention (FFDC) mechanism. FFDC employs a lightweight verifier to dynamically adjust action chunk lengths and incorporates Mixture-of-Horizon Training to enhance coverage of long-horizon trajectories. The approach integrates predicted actions, visual dynamics, real observations, and language instructions, leveraging causal attention for consistency evaluation. Experiments demonstrate that the method reduces forward computation by 69.10%, shortens execution time by 34.02%, and improves success rates by 2.54% on RoboTwin, with a 35% absolute gain in real-robot trials.

0 citationsRead paper

CLAP: Contrastive Latent Action Pretraining for Learning Vision-Language-Action Models from Human Videos

Jan 07, 2026arXiv.org

This work addresses the limitations of current vision-language-action (VLA) models, which are hindered by the scarcity of robotic demonstration data and the susceptibility of human-video-based latent action approaches to visual distractions, impeding the extraction of executable skills. To overcome these challenges, we propose the Contrastive Latent Action Pretraining (CLAP) framework, which aligns the visual latent space of human videos with robot proprioceptive trajectories through contrastive learning and maps actions to an executable quantized codebook. We introduce a dual-branch VLA architecture—comprising CLAP-NTP and CLAP-RF—that integrates Rectified Flow with knowledge-matching regularization to effectively mitigate catastrophic forgetting during fine-tuning. Experiments demonstrate that our approach significantly outperforms baseline methods in skill transfer tasks, achieving superior instruction following, object generalization, and high-frequency precise manipulation capabilities.

1 citationsRead paper