🤖 AI Summary
This work addresses the sensitivity of vision-language-action (VLA) policies to initial arm configurations in bimanual humanoid manipulation, which often leads to task failure or incorrect hand selection. The study introduces the novel concept of “policy-induced hand prior” and systematically quantifies the causal influence of initial arm poses on hand preference, revealing a critical link between training data distribution and pose robustness. By integrating HandPriorScore, residual hand bias, and target responsiveness metrics with multi-policy evaluation and wrist-camera observation analysis, the authors identify low-robustness configurations. They then expand the training pose coverage and apply targeted data augmentation. Experiments across 17 initial configurations demonstrate strong policy-pose interaction effects, significantly improving overall task success rates—particularly for previously underperforming poses.
📝 Abstract
Vision-language-action (VLA) policies are expected to operate robustly across variations in the robot's initial configuration, yet aggregate task success can conceal pose-specific failures and inappropriate hand selection. This work investigates initial-pose dependence in VLA-based humanoid dual-arm manipulation. We characterize the initial-condition-dependent early hand preference as a policy-induced hand prior and quantify it using HandPriorScore, residual hand bias, and target responsiveness. Evaluations across multiple policies and 17 initial configurations reveal strong initial-pose--policy interactions: the same pose produces substantially different success rates across policies, while a single policy exhibits large performance variation across poses. Specific initial arm configurations can suppress or induce an asymmetric hand preference, with the resulting effect varying in direction and strength across policies. Wrist-camera observations also influence hand selection and task performance. Expanding initial-pose coverage in the training dataset substantially improves robustness, while targeted augmentation around a low-performing configuration increases its success rate. Comparisons across training configurations show that sufficient exposure to the target simulation task is beneficial, whereas the effect of real or auxiliary data depends on pose coverage, simulation ratio, and observation availability. These findings characterize a pose-conditioned hand prior, identify a localized initial arm configuration as a causal handle on hand-selection behavior, and demonstrate how data coverage and training composition affect initial-pose robustness.