DistAL: Distance-based Advantage Learning for VLA Fine-Tuning

📅 2026-09-16
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为提高VLA模型的性能,通过引入基于距离的优势学习方法DistAL,使用嵌入空间距离作为奖励,生成更优质的价值函数,从而提升下游任务成功率。
📝 Abstract
Vision-language-action models (VLAs) have trans- formed the field of robotic manipulation in recent years by combining the semantic understanding of LLMs with the precise control of flow-matching policies. Advantage conditioning is a recent technique that iteratively improves VLAs by training a value function on deployment data and using this to train an advantage-conditioned policy. Previous works have only applied simple, low-information success/failure rewards, which leave the value function unable to distinguish states of differing quality beyond how far along the task they appear. Motivated by an exploration of out-of-distribution (OOD) detection methods, we introduce Distance-based Advantage Learning (DistAL), which, by using an embedding space distance as a reward, produces a more informative value function and subsequently a higher downstream task success rate. We validate our method on a series of simulation benchmarks and dexterous bi-manual manipulation tasks on real hardware.
Problem

Research questions and friction points this paper is trying to address.

Advantage Learning
Value Function
Reward
OOD Detection
Embedding Space
Innovation

Methods, ideas, or system contributions that make the work stand out.

Distance-based Advantage Learning
embedding space distance
value function
downstream task success rate
🔎 Similar Papers
2024-08-10AAAI Conference on Artificial IntelligenceCitations: 30
💼 Related Jobs
No related jobs found.