Quantile-Coupled Flow Matching for Distributional Reinforcement Learning
This work addresses a critical inconsistency in existing conditional flow matching (CFM) approaches for distributional reinforcement learning, where arbitrary source–target pairings yield losses misaligned with the Wasserstein distance, thereby violating the contraction property of the Bellman operator. To resolve this, the authors propose FlowIQN, which constructs quantile-aligned, monotonic optimal transport couplings by sorting source samples and Bellman targets within each minibatch, ensuring flow trajectories consistent with the Wasserstein metric. FlowIQN provides the first explicit Wasserstein projection guarantee for flow-matching distributional critics and incorporates a shortcut inference model to enhance computational efficiency. Empirical results demonstrate that FlowIQN significantly improves the Wasserstein accuracy of return distributions and achieves strong performance across multiple offline reinforcement learning benchmarks under various policy extraction settings, offering both theoretical rigor and practical effectiveness.