Generative Modeling by Minimizing the Wasserstein-2 Loss
This work addresses the challenge of minimizing the Wasserstein-2 (W₂) distance in unsupervised generative modeling. We propose the first explicit, distribution-dependent ordinary differential equation (ODE) characterizing the W₂ gradient flow, and theoretically prove that its time-marginal distributions converge rigorously to the target data distribution. Methodologically, we design a persistent Euler discretization algorithm that couples with the gradient flow structure, circumventing the instability inherent in conventional adversarial training. Our key contributions are: (1) the first explicit ODE formulation of the W₂ gradient flow; and (2) a novel persistence training mechanism that enhances discretization fidelity and convergence robustness. Experiments demonstrate that our approach significantly outperforms WGAN on both high- and low-dimensional benchmarks; moreover, increasing persistence strength further improves generation quality and training stability.