Iterative Distillation for Reward-Guided Fine-Tuning of Diffusion Models in Biomolecular Design
In biomolecular design, reward functions are often non-differentiable—e.g., derived from physics-based simulations or domain-specific heuristics—posing challenges for reward-guided diffusion model training, including instability, low sample efficiency, and mode collapse. To address this, we propose an iterative policy distillation framework that enables efficient fine-tuning of diffusion models under arbitrary (including non-differentiable) rewards. Our method leverages off-policy data reuse, soft-optimal policy modeling, and KL-divergence-regularized policy updates to stabilize learning and preserve diversity. Compared to existing reinforcement learning–based approaches, it significantly improves training stability and sample efficiency while avoiding mode collapse. We validate the framework across protein, small-molecule, and regulatory DNA design tasks, achieving state-of-the-art reward optimization performance in all settings. Results demonstrate the method’s generality, robustness to diverse reward formulations, and scientific validity in real-world biomolecular design applications.