Learning Multimodal One-step Flow Policy via Value-weighted Optimal Transport

📅 2026-09-14
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决离线强化学习中多模态动作分布的表示问题,提出OptiFlow方法,通过最优传输指导单步流策略学习,避免模式坍塌和过估计偏差。
📝 Abstract
Offline reinforcement learning aims to learn a policy solely from fixed datasets, which often contain multimodal action distributions. Flow policies can naturally represent such multimodal behaviors, but learning an efficient one-step flow policy remains challenging: standard value guidance often leads to mode collapse or exploits overestimation bias in out-of-distribution regions. To address this, we introduce One-step Flow policy via Optimal Transport (OptiFlow), a framework for one-step flow policy learning as a structured sample-allocation problem. OptiFlow jointly trains a value-aware reference flow policy and an efficient one-step policy, coupling their action samples through state-wise entropic optimal transport. For each state, critic-estimated values define the priority of distillation target actions, while the action-distance cost ensures geometrically compatible pairings. By avoiding direct critic maximization, our transport-guided approach enables in-distribution exploitation by anchoring the one-step policy to high-value, dataset-supported modes without the risk of out-of-distribution divergence. Experimental results demonstrate that OptiFlow effectively captures optimal multimodal behaviors and achieves strong performance across diverse offline RL benchmarks. Our code is available at https://github.com/Yonsei-DILLab/OptiFlow.
Problem

Research questions and friction points this paper is trying to address.

offline reinforcement learning
multimodal action distributions
mode collapse
overestimation bias
Innovation

Methods, ideas, or system contributions that make the work stand out.

Optimal Transport
One-step Flow Policy
Value-aware Reference
Structured Sample-allocation
Entropic Optimal Transport
💼 Related Jobs
No related jobs found.
J
Jaehun Shon
Department of Artificial Intelligence, Yonsei University
J
Jinha Choi
Department of Artificial Intelligence, Yonsei University
J
Jongwook Jeon
Department of Artificial Intelligence, Yonsei University
Jongmin Lee
Jongmin Lee
Yonsei University
Machine LearningReinforcement Learning