OrthoSkillVLA: Continual Skill Learning via Gradient-Informed Skill Subspace Adaptation

📅 2026-08-19
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
该研究提出OrthoSkillVLA框架,通过为预训练的视觉-语言-动作模型中的不同组件设置独立的子空间约束及引入轻量级特征感知MoE解码器,解决了连续技能学习中的灾难性遗忘问题。
📝 Abstract
Pretrained Vision-Language-Action models provide a strong foundation for robot learning, but sequentially adapting them to diverse skills can perturb the representations and velocity mappings used by previous skills, leading to catastrophic forgetting. Architecture-based approaches improve retention by isolating skills but lead to increased inference footprint. Recent subspace-constrained methods restrict parameter updates in an orthogonal subspace to minimize interference but impose a unified constraint on the entire model. We analyze the distinct roles of internal VLA components and identify two VLA-specific challenges. First, the VLM maintains broad semantic representations, making it vulnerable to capacity exhaustion, whereas the ActionHead refines semantics into localized velocity patterns that are highly sensitive to perturbations. Second, the final velocity decoder serves as a readout layer. Freezing it forms an output-stage expressivity bottleneck, while updating it risks overwriting previous velocity mappings. To this end, we propose OrthoSkillVLA, a parameter-efficient framework for continual skill learning in pretrained VLA models without demonstration replay. Given the representation heterogeneity, we impose separate subspace constraints on the VLM and ActionHead, preserving reusable semantic capacity while protecting localized velocity patterns. For the output layer, we introduce a lightweight feature-aware MoE decoder, where each skill is allocated a compact expert and a training-free router selects the expert according to feature-space affinity. Extensive simulated and real-world evaluations, together with ablations, demonstrate that OrthoSkillVLA better preserves prior skills while acquiring new ones.
Problem

Research questions and friction points this paper is trying to address.

continual skill learning
catastrophic forgetting
subspace constraints
pretrained VLA models
Innovation

Methods, ideas, or system contributions that make the work stand out.

Gradient-Informed Skill Subspace Adaptation
Subspace Constraints
Feature-Aware MoE Decoder
Continual Skill Learning
🔎 Similar Papers
No similar papers found.
Jiaqi Wang
Jiaqi Wang
Harbin Institute of Technology Shenzhen & Pengcheng Laboratory, Computer Science
Spiking Neural NetworkBrain DecodingSpeechBrain Computer Interface
Z
Zhou Fang
School of Computer Science and Engineering, Southeast University, China; Key Laboratory of New Generation Artificial Intelligence Technology and Its Interdisciplinary Applications, Ministry of Education, China
Qiongfeng Shi
Qiongfeng Shi
Southeast University; National University of Singapore
Flexible electronicsSensorsEnergy harvestersIntelligent systems
Y
Yi Zhou
School of Computer Science and Engineering, Southeast University, China; Key Laboratory of New Generation Artificial Intelligence Technology and Its Interdisciplinary Applications, Ministry of Education, China