VGFM: Expressive Robot Policies via Dense Value Guidance in Flow Matching

📅 2026-09-12
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出VGFM方法,通过在流匹配模型中引入密集的价值引导来改进机器人策略学习,避免了反向传播和额外的算法开销,适用于复杂的离线强化学习任务。
📝 Abstract
Recent robot learning paradigms increasingly rely on large offline datasets of robotic interactions to train control policies. Expressive generative models enable rich and multimodal action representations, expanding the capability of this paradigm for complex robotic control. However, policy improvement with multi-step generative actors remains challenging. In offline reinforcement learning (RL), incorporating value-based objectives along generative trajectories often introduces substantial training complexity, including backpropagation through time (BPTT), auxiliary architectures, or distillation losses. We propose Value-Guided Flow Matching (VGFM), a scalable offline RL framework that enables dense value-guided shaping within a flow-based policy while avoiding BPTT and additional algorithmic overhead. VGFM parameterizes the policy as a conditional flow-matching model in action (x-prediction) space, ensuring that each intermediate flow step produces a valid robot action that can be directly evaluated by a standard offline RL critic. This design allows value guidance to be applied at randomly sampled flow times without differentiating through the entire generative trajectory, while preserving inference-time flexibility by varying the discretization of the underlying flow ODE without retraining. Evaluated on robotic locomotion and manipulation tasks in OGBench, VGFM achieves strong performance across a wide range of tasks under rigorous evaluation protocols. With minimal hyperparameter tuning, these results demonstrate that VGFM provides a simple, scalable, and effective approach for expressive policy learning in long-horizon, goal-oriented robotic control.
Problem

Research questions and friction points this paper is trying to address.

offline reinforcement learning
value-based objectives
generative models
policy improvement
training complexity
Innovation

Methods, ideas, or system contributions that make the work stand out.

Value-Guided Flow Matching
offline reinforcement learning
flow-based policy
dense value guidance
robotic control
💼 Related Jobs
No related jobs found.
P
Prajwal Koirala
Sibley School of Mechanical and Aerospace Engineering, Cornell University
Mark Campbell
Mark Campbell
John A. Mellowes ‘60 Professor, Cornell University
Autonomy: EstimationControlAerospaceRoboticsSpace Systems