Harness VLA: Steering Frozen VLAs into Reliable Manipulation Primitives via Memory-Guided Agents
Existing vision-language-action (VLA) models exhibit limited generalization under perturbations such as semantic re-targeting, object re-binding, layout changes, and unstable contacts, while purely analytical primitives struggle with irregular grasps and complex interactions. This work proposes a memory-augmented agent framework that decouples a frozen VLA—used as a retry-capable primitive for contact-intensive manipulation—from a small set of fixed analytical primitives. A memory-guided planner handles non-contact phases and semantic re-localization, invoking the VLA only during localized contact stages. The system learns the applicability boundaries of each primitive through execution trajectories, success heuristics, and failure models, thereby extending the VLA’s capabilities without fine-tuning. The approach outperforms the strongest baseline by 38.6 and 25.4 percentage points on LIBERO-Pro and RoboCasa365, respectively, and achieves a 58.4% success rate on RoboTwin C2R.