Not All Layers Need Tuning: Diagnosing and Directing Adaptation in Vision-Language-Action Models

📅 2026-09-15
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文针对视觉-语言-动作模型在新环境下的微调成本问题,提出一种诊断并定向适应的方法,通过估计各区域适应成本并分配适应器来优化微调过程。
📝 Abstract
Fine-tuning a Vision-Language-Action (VLA) model for a new deployment environment is expensive, yet most methods apply uniform-capacity adapters to every network region as if every region requires equal adjustment. This paper tests that assumption on five architecturally diverse VLAs (OpenVLA-OFT, $π_0$, SmolVLA, DTP, Octo; 93M-7B parameters). Measuring per-region adaptation cost as normalized parameter displacement under region-isolated fine-tuning reveals an adaptation spectrum in which appearance shifts concentrate cost in the vision encoder, instruction shifts in the language backbone, and novel-object shifts in the vision encoder together with the action head, across all five architectures. To exploit this structure, we introduce a pipeline that observes, diagnoses, allocates, and adapts. From ten unlabeled target observations and without fine-tuning, the diagnostic estimates per-region cost by combining reference-free gradient and Monte Carlo Dropout signals with a Centered Kernel Alignment score against a cached source reference; the allocator converts the estimates into variable-rank LoRA adapters under a parameter budget and freezes well-calibrated regions; and standard LoRA fine-tuning trains the resulting adapters. The diagnostic ranks regions within each deployment at a median Spearman of 0.91, and the allocation matches or exceeds uniform LoRA at every budget we tested on LIBERO and CALVIN. On a physical xArm-7, the pipeline matches full fine-tuning under an instruction-wording shift with 0.04% of its trainable parameters, and on five held-out scenes evaluated without retraining it leads every baseline, with 11-23 successes of 30 rollouts against 8-18 for the strongest parameter-efficient baseline at equal or larger budgets and 2-11 for full fine-tuning. These results suggest that adaptation cost in VLAs is structured enough to measure before fine-tuning begins.
Problem

Research questions and friction points this paper is trying to address.

Vision-Language-Action
Fine-tuning
Adaptation Cost
Innovation

Methods, ideas, or system contributions that make the work stand out.

Region-specific Adaptation
Low-Rank Adapters (LoRA)
Parameter Efficiency
Pre-Fine-Tuning Diagnosis
Cost Estimation
S
Shahram Najam Syed
Robotics Institute, Carnegie Mellon University, Pittsburgh, USA
A
Arthur Jakobsson
Robotics Institute, Carnegie Mellon University, Pittsburgh, USA
P
Prayuj Sachdev
Department of Electrical and Computer Engineering, Carnegie Mellon University, Pittsburgh, USA
Jeffrey Ichnowski
Jeffrey Ichnowski
Carnegie Mellon University
RoboticsManipulationMotion Planning