Training-Free Action Correction for VLA Model Failures via Language Feedback

📅 2026-08-30
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出CorrectVLA框架,通过语言反馈调整动作幅度,无需重新训练即可修正VLA模型在部署中的执行偏差问题。
📝 Abstract
Vision-Language-Action (VLA) models demonstrate strong semantic understanding yet exhibit systematic failures during deployment. The conditions under which these failures occur, and whether they can be corrected without retraining, remain poorly understood. In this paper, we take steps toward addressing this gap. We present CorrectVLA, a framework that translates task-level natural language corrections into additive action magnitude adjustments without modifying policy weights. A human provides a single task-level correction, applied uniformly across all rollouts without per-episode intervention. In simulation, CorrectVLA recovers execution misalignment failures across both in-distribution and OOD tasks. In real-robot experiments on a UFactory xArm7 under environment shift, CorrectVLA restores near-perfect success where the base policy almost entirely breaks down, generalizing across object locations and identities. Through a taxonomy of failure modes on LIBERO-90, we find that execution misalignment failures, where the policy reaches the correct target but miscalibrates action magnitudes, represent the correctable subset, while other failure modes where semantic comprehension itself breaks down are not amenable to this approach. The approach succeeds when policies possess strategic correctness and fails when fundamental comprehension is absent, establishing a practical operational boundary for inference-time correction.
Problem

Research questions and friction points this paper is trying to address.

Vision-Language-Action Models
Systematic Failures
Language Feedback
Action Correction
Deployment
Innovation

Methods, ideas, or system contributions that make the work stand out.

Training-Free Action Correction
Language Feedback
Execution Misalignment Failures
Additive Action Magnitude Adjustments
Inference-Time Correction
O
Owen Kwon
Biomedical Engineering, Carnegie Mellon University, Pittsburgh PA 15213, USA
P
Pablo Ortega-Kral
Robotics Institute, Carnegie Mellon University, Pittsburgh PA 15213, USA
A
Arthur Bucker
Robotics Institute, Carnegie Mellon University, Pittsburgh PA 15213, USA
Jean Oh
Jean Oh
Robotics Institute, Carnegie Mellon University
RoboticsMultimodal PerceptionSocial NavigationLanguage-Vision intersectionArtificial Intelligence