Unlocking Lossless Speedups in LLMs via Discrete Diffusion

📅 2026-09-03
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决大型语言模型自回归结构导致的慢速顺序生成问题,本文提出通过扩散增强的语言模型,利用扩散并行生成多个令牌,实现加速且不损失质量。
📝 Abstract
Large Language Models (LLMs) owe much of their success to next-token prediction (NTP), but their autoregressive (AR) structure requires slow, sequential token generation. To overcome this bottleneck, we introduce diffusion-augmented LLMs, a new class of models that defines an AR model distribution while using diffusion to draw multiple tokens in parallel from that distribution. We decouple the parameters of these models into two sets: AR weights, trained using the standard NTP objective, and lightweight diffusion weights, trained to generate multiple tokens simultaneously. The diffusion weights are learned through a simple Diffusion Distillation phase that adds negligible overhead to existing LLM training pipelines. We also introduce $Ψ$-Spec, a family of samplers that enables lossless acceleration and inference-time scaling at a fixed context length. Unlike speculative decoding, our method requires no separate draft model. Unlike diffusion LLMs (d-LLMs), it accelerates generation without sacrificing the quality of the underlying AR model. The resulting models, called Uno, can be trained from scratch or built by augmenting existing open-weight AR LLMs. Uno achieves higher throughput than leading speculative-decoding methods at every evaluated batch size and delivers up to $3\times$ speedups over the base AR model, including at the largest batch size supported by the device. Notably, our 8B Uno model outperforms the leading open d-LLM, the 26B DiffusionGemma, and the proprietary Mercury 2 across all evaluated benchmarks in agentic tool use, coding, and long-context reasoning. We release code and checkpoints at: https://s-sahoo.github.io/uno/
Problem

Research questions and friction points this paper is trying to address.

Large Language Models
autoregressive structure
token generation
speed bottleneck
Innovation

Methods, ideas, or system contributions that make the work stand out.

Diffusion-Augmented LLMs
Parallel Token Generation
Ψ-Spec Samplers
Lossless Acceleration
Inference-Time Scaling
Subham Sekhar Sahoo
Subham Sekhar Sahoo
Cornell University
Diffusion Language ModelsGenerative AI
Lingjie Chen
Lingjie Chen
Computer Science, University of Illinois Urbana-Champaign
Trustworth LLMMechanistic Interpretability
Khiem Pham
Khiem Pham
PhD student
machine learning
Jonathan Geuter
Jonathan Geuter
PhD student, Harvard University
Machine LearningOptimal TransportGenerative ModellingReinforcement Learning
C
Chaitanya Dwivedi
Institue of Foundation Models
V
Varad Pimpalkhute
Institue of Foundation Models
Yash Akhauri
Yash Akhauri
Google Research, PhD Candidate at Cornell University
Computer VisionQuantized Deep LearningAutoMLIntelligent Systems
Alexander Moreno
Alexander Moreno
Institute of Foundation Models, MBZUAI
LLM pre-trainingtraining dynamicsfoundation models
Mikhail Yurochkin
Mikhail Yurochkin
Staff AI Scientist, IFM MBZUAI, ex MIT-IBM Watson AI Lab
Machine LearningFoundation ModelsEvaluationModel Fusion
Zhenting Wang
Zhenting Wang
Accenture; Rutgers University
Mostafa Elhoushi
Mostafa Elhoushi
Research Scientist @ Cerebras Systems
Deep LearningMachine LearningSensorsQuantum Computing
Nolan Dey
Nolan Dey
Cerebras Systems
Large language modelsTraining efficiencySparsityExplainable AI
Shane Bergsma
Shane Bergsma
Cerebras Systems
Machine LearningArtificial IntelligenceNatural Language Processing
Joel Hestness
Joel Hestness
Distinguished Research Scientist, Cerebras Systems
Deep LearningLanguage UnderstandingHeterogeneous SystemsHigh-performance Computing
John Thickstun
John Thickstun
Assistant Professor, Cornell University
Machine LearningGenerative ModelsMusic TechnologyNatural Language Processing
Eric Xing
Eric Xing
President at Mohamed bin Zayed University of AI, Professor of Computer Science, Carnegie Mellon U
Machine LearningML SystemsStatisticsNetwork AnalysisAI4Science
Zhengzhong Liu
Zhengzhong Liu
Institute of Foundation Models
Natural Language ProcessingMachine Learning