Robust Bimanual Vision-Language-Action Models via Embarrassingly Simple Modality Masking

📅 2026-08-23
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文针对双臂机器人操作中动作不连贯和执行失败的问题,提出了一种简单的模态掩蔽机制(M3),通过在训练时随机掩蔽部分模态通道来提高模型的鲁棒性。
📝 Abstract
Query-based Vision-Language-Action (VLA) models offer low-latency inference that is attractive for bimanual robotic manipulation, but we observe that they can still exhibit discontinuous actions and execution failures in complex dual-arm tasks. We hypothesize that unstable multi-view and language fusion is one contributing factor in these failures, often coinciding with attention spreading to distracting regions. To improve robustness, we introduce the Modality Masking Mechanism (M3), an embarrassingly simple, training-only strategy that requires no architectural changes or large-scale robot pretraining. M3 stochastically masks subsets of modality channels during training, exposing the policy to controlled partial observations and encouraging it to rely less on distracting cues and more on evidence that remains reliable. We evaluate M3 on ten bimanual tasks from RoboTwin 2.0 and on three long-horizon real-world tasks. Compared with the Adapter baseline, M3 improves average success by 21.7% in the Clean setting and 11.4% in Clean2Rand, where policies are trained on clean demonstrations and evaluated on randomized scenes, while also improving averaged real-world full-task success by over 30%. These results suggest that structured training-time masking is a practical way to improve the robustness of query-based VLA policies for bimanual manipulation.
Problem

Research questions and friction points this paper is trying to address.

Vision-Language-Action
bimanual manipulation
discontinuous actions
execution failures
multi-view and language fusion
Innovation

Methods, ideas, or system contributions that make the work stand out.

Modality Masking Mechanism
bimanual manipulation
Vision-Language-Action models
robustness improvement
🔎 Similar Papers
💼 Related Jobs
No related jobs found.