Vision-Language Models Unlock Task-Centric Latent Actions
This work addresses the vulnerability of existing latent action models to task-irrelevant distractors, which often leads to the erroneous encoding of noise as action signals. To mitigate this, the authors propose a novel approach that leverages the commonsense reasoning capabilities of vision-language models (VLMs) to generate task-aware representations distinguishing controllable dynamics from noise. Specifically, task-oriented natural language prompts—such as “ignore distractors”—are used to guide VLMs in producing supervision signals that, in an unsupervised setting, steer latent action models toward learning task-centric action representations. Evaluated on the Distracting MetaWorld benchmark, the method improves downstream task success rates by up to sixfold, significantly enhances action semantic consistency, and effectively suppresses interference. The study also reveals notable differences in prompt sensitivity and performance across various VLMs in the context of action representation learning.