Direct Ascent Synthesis: Revealing Hidden Generative Capabilities in Discriminative Models
This work challenges the conventional dichotomy between discriminative and generative models, providing the first theoretical proof that standard discriminative models—such as CLIP—implicitly encode rich generative knowledge. To harness this, we propose Direct Ascent Synthesis (DAS), a training-free, multi-scale latent-space optimization method: it performs concurrent gradient ascent across resolution levels (1×1 to 224×224), integrating hierarchical latent inversion with natural-image 1/f² spectral priors. DAS enables zero-shot text-to-image generation and cross-domain style transfer without any model fine-tuning or adversarial training. Remarkably, it achieves image fidelity comparable to dedicated generative models while markedly suppressing artifacts and rigorously preserving human visual statistics—e.g., natural-scene spectral decay and structural coherence. By eliminating reliance on explicit generative architectures or adversarial objectives, DAS redefines the boundary between discriminative and generative modeling, demonstrating that high-fidelity synthesis can emerge directly from off-the-shelf discriminative representations.