Institution profile

Answer AI

Industry researchnorthamerica · us
Official website
Research library2linked papers
Opportunities0open roles
Selected work

Representative Papers

Direct Ascent Synthesis: Revealing Hidden Generative Capabilities in Discriminative Models

Feb 11, 2025

This work challenges the conventional dichotomy between discriminative and generative models, providing the first theoretical proof that standard discriminative models—such as CLIP—implicitly encode rich generative knowledge. To harness this, we propose Direct Ascent Synthesis (DAS), a training-free, multi-scale latent-space optimization method: it performs concurrent gradient ascent across resolution levels (1×1 to 224×224), integrating hierarchical latent inversion with natural-image 1/f² spectral priors. DAS enables zero-shot text-to-image generation and cross-domain style transfer without any model fine-tuning or adversarial training. Remarkably, it achieves image fidelity comparable to dedicated generative models while markedly suppressing artifacts and rigorously preserving human visual statistics—e.g., natural-scene spectral decay and structural coherence. By eliminating reliance on explicit generative architectures or adversarial objectives, DAS redefines the boundary between discriminative and generative modeling, demonstrating that high-fidelity synthesis can emerge directly from off-the-shelf discriminative representations.

0 citationsRead paper

It's All in The [MASK]: Simple Instruction-Tuning Enables BERT-like Masked Language Models As Generative Classifiers

Feb 06, 2025

This work addresses the limited generalization of encoder-only models (e.g., BERT, ModernBERT) on generative classification tasks, aiming to match decoder-based large language models (LLMs) without task-specific classification heads. Methodologically, it pioneers systematic reuse of the standard masked language modeling (MLM) head—combined with lightweight instruction tuning—to perform zero-shot and fine-tuned generative classification directly via the [MASK] token. Key contributions are threefold: (1) first theoretical and empirical validation that the MLM head can effectively substitute conventional classification heads; (2) identification of modern, diverse pretraining data as a critical prerequisite for unlocking this capability; and (3) demonstration that ModernBERT-Large-Instruct achieves 93% of Llama3-1B’s MMLU score with 60% fewer parameters, surpassing same-scale LLMs in zero-shot accuracy and outperforming traditional classification-head paradigms after fine-tuning.

0 citationsRead paper
Recent publications

Latest Papers

Direct Ascent Synthesis: Revealing Hidden Generative Capabilities in Discriminative Models

Feb 11, 2025

This work challenges the conventional dichotomy between discriminative and generative models, providing the first theoretical proof that standard discriminative models—such as CLIP—implicitly encode rich generative knowledge. To harness this, we propose Direct Ascent Synthesis (DAS), a training-free, multi-scale latent-space optimization method: it performs concurrent gradient ascent across resolution levels (1×1 to 224×224), integrating hierarchical latent inversion with natural-image 1/f² spectral priors. DAS enables zero-shot text-to-image generation and cross-domain style transfer without any model fine-tuning or adversarial training. Remarkably, it achieves image fidelity comparable to dedicated generative models while markedly suppressing artifacts and rigorously preserving human visual statistics—e.g., natural-scene spectral decay and structural coherence. By eliminating reliance on explicit generative architectures or adversarial objectives, DAS redefines the boundary between discriminative and generative modeling, demonstrating that high-fidelity synthesis can emerge directly from off-the-shelf discriminative representations.

0 citationsRead paper

It's All in The [MASK]: Simple Instruction-Tuning Enables BERT-like Masked Language Models As Generative Classifiers

Feb 06, 2025

This work addresses the limited generalization of encoder-only models (e.g., BERT, ModernBERT) on generative classification tasks, aiming to match decoder-based large language models (LLMs) without task-specific classification heads. Methodologically, it pioneers systematic reuse of the standard masked language modeling (MLM) head—combined with lightweight instruction tuning—to perform zero-shot and fine-tuned generative classification directly via the [MASK] token. Key contributions are threefold: (1) first theoretical and empirical validation that the MLM head can effectively substitute conventional classification heads; (2) identification of modern, diverse pretraining data as a critical prerequisite for unlocking this capability; and (3) demonstration that ModernBERT-Large-Instruct achieves 93% of Llama3-1B’s MMLU score with 60% fewer parameters, surpassing same-scale LLMs in zero-shot accuracy and outperforming traditional classification-head paradigms after fine-tuning.

0 citationsRead paper