Generalized Discrete Diffusion with Self-Correction
Existing self-correction methods for discrete diffusion models are typically introduced during inference or post-training, exhibiting limited generalization and often degrading performance. This work proposes the Self-Correcting Discrete Diffusion (SCDD) model, which, for the first time, integrates an explicit state-transition-based self-correction mechanism directly into pretraining within a discrete-time framework. SCDD eliminates redundant re-masking steps and relies solely on a uniform absorption objective for learning. By simplifying the noise schedule and combining BERT-style pretraining with parallel decoding, the method substantially improves decoding efficiency—demonstrated on GPT-2-scale experiments—while preserving high generation quality.