Self-Corrected Flow Distillation for Consistent One-Step and Few-Step Text-to-Image Generation
Existing flow-matching models require numerous function evaluations during sampling, compromising the trade-off between efficiency and generation quality—particularly yielding poor consistency in single-step or few-step sampling. This paper proposes a self-correcting flow distillation framework that, for the first time, jointly integrates consistency modeling and adversarial training into the flow-matching paradigm. Leveraging knowledge distillation, our approach enables high-fidelity, highly consistent one-step and few-step text-to-image synthesis. Crucially, it preserves sampling efficiency while substantially improving generation fidelity and stability. Quantitative and qualitative evaluations on CelebA-HQ demonstrate superior performance over state-of-the-art methods. Moreover, zero-shot evaluation on COCO shows significant improvements in text-image alignment and fine-grained detail preservation. The implementation is publicly available.