GFlowNet Training by Policy Gradients
GFlowNets suffer from low training efficiency and unstable gradient estimation in combinatorial object generation due to strict flow conservation constraints. To address this, we propose the first policy-gradient-based GFlowNet training framework. Our method reformulates flow conservation as a policy optimization objective, enabling joint training of forward and backward policies without explicit flow matching. We provide theoretical convergence guarantees and introduce a coupled update mechanism to reduce gradient variance. Experiments across multiple synthetic and real-world datasets demonstrate that our approach significantly improves sample quality, training stability, and fidelity to the target distribution—particularly under sparse-reward settings, where it exhibits superior robustness.