SOP: A Scalable Online Post-Training System for Vision-Language-Action Models
This work proposes SOP, the first online, multi-robot collaborative, and multi-task post-training framework for general-purpose vision-language-action (VLA) models. Existing VLA post-training methods are typically offline, single-machine, or task-specific, limiting their capacity for efficient online adaptation and large-scale real-world learning. SOP addresses this gap through a closed-loop bitstream architecture that tightly couples a fleet of robots with a cloud-based learner. The system integrates interactive imitation learning (HG-DAgger) and reinforcement learning (RECAP), enabling asynchronous policy updates and human-in-the-loop interventions. Evaluated on real-world tasks such as cloth folding and box assembly, SOP significantly improves pretrained model performance within hours, with gains scaling nearly linearly with the number of robots while preserving the generality of a single shared policy.