π€ AI Summary
Existing diffusion models struggle to achieve effective and theoretically guaranteed control during inference when guided by distribution-level rewards, such as diversity or population statistics. This work addresses this challenge by formulating it as the approximation of a tilted measure within a mean-field framework and introduces a weighted interacting particle scheme that, for the first time, provides theoretical guarantees for distribution-level reward control. The proposed approach unifies existing pointwise reward control paradigms and establishes a theoretical foundation for batch guidance strategies. Empirical results demonstrate that the method accurately approximates target distributions in low-dimensional analytically tractable tasks and achieves superior performance in high-dimensional protein conformation generation.
π Abstract
Diffusion models are increasingly used as controllable samplers, whose generations can be steered at inference time according to a chosen reward function. While such rewards are typically defined on individual samples, for many applications it is desirable to steer according to distribution-level rewards, for example to calibrate with population-level information or to encourage diversity. In both cases, simply incorporating the reward gradient into the dynamics, while often effective, comes with few theoretical guarantees on the sampled distribution. For pointwise rewards, recent work has therefore sought to develop a principled framework for targeting a prescribed tilted distribution using particle reweighting. However, an analogous theoretically-grounded approach for distributional rewards is currently lacking. In this work, we formulate inference-time distributional control as targeting a tilted measure under a mean-field framework, and derive a weighted interacting particle scheme to target it in a principled manner. Our framework recovers pointwise-reward steering as a special case, while providing a theoretical foundation for existing batch-level steering methods. Empirically, we verify that the procedure correctly targets the prescribed distribution in tractable low-dimensional settings, and investigate its behaviour in higher-dimensional protein conformation tasks.