LLaDA-Image: Building Strong Image Generators with Fully Open Training Recipes

📅 2026-09-03
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出LLaDA-Image框架,通过仅图像预训练和中程训练建立强大的视觉生成先验,并结合优化技术生成高保真度图像,同时在基准测试上取得领先成绩。
📝 Abstract
We introduce LLaDA-Image, a unified framework that pairs a 6B Diffusion Transformer (DiT) trained from scratch with a frozen vision-language understanding module built on the LLaDA2.0-Mini diffusion language model backbone. Instead of relying heavily on paired image-text data from the beginning, we first build a strong visual generative prior through image-only pre-training and mid-training. The generation pipeline comprises 220M samples, 98 of which are real images. For efficient and scalable optimization, we use parameter-free RMSNorm throughout the DiT together with the Muon optimizer. The resulting unified model produces highly photorealistic images while accurately following fine-grained editing instructions. We further distill LLaDA-Image into LLaDA-Image-Turbo, enabling fast inference in 2-4 sampling steps. On Qwen-Image-Bench, LLaDA-Image achieves overall scores of 53.53 and 53.38 on the English and Chinese tracks, respectively, setting a new state-of-the-art among open-source models on both tracks. To support further research on capable and efficient generative models, we release our model weights, training code, and detailed recipes.
Problem

Research questions and friction points this paper is trying to address.

Image Generation
Diffusion Transformer
Vision-Language Understanding
Fine-Grained Editing
Innovation

Methods, ideas, or system contributions that make the work stand out.

Diffusion Transformer
vision-language understanding
image-only pre-training
RMSNorm
fast inference
🔎 Similar Papers
No similar papers found.
C
Chuyan Chen
AGI Research Center, Inclusion AI
H
Haoxing Chen
AGI Research Center, Inclusion AI
K
Kun Chen
AGI Research Center, Inclusion AI
Zhenglin Cheng
Zhenglin Cheng
Zhejiang University & Westlake University, SII
Multimodal LearningDiffusion Models
L
Long Cui
AGI Research Center, Inclusion AI
R
Ruishan Fang
AGI Research Center, Inclusion AI
Zhangxuan Gu
Zhangxuan Gu
Ant Group
computer vision
Z
Zhicheng Huang
AGI Research Center, Inclusion AI
Zhenzhong Lan
Zhenzhong Lan
School of Engineering, Westlake University
NLPComputer VisionMultimedia
Y
Yuanting Lei
AGI Research Center, Inclusion AI
H
Haoquan Li
AGI Research Center, Inclusion AI
Jianguo Li
Jianguo Li
Director, Ant Group
deep learningcomputer visionmachine learningsystem
R
Rongchuan Li
AGI Research Center, Inclusion AI
S
Sidu Li
AGI Research Center, Inclusion AI
T
Tao Lin
AGI Research Center, Inclusion AI
D
Deyuan Liu
AGI Research Center, Inclusion AI
J
Jiacheng Liu
AGI Research Center, Inclusion AI
L
Lin Liu
AGI Research Center, Inclusion AI
Y
Yuxuan Lou
AGI Research Center, Inclusion AI
Z
Zhisheng Lu
AGI Research Center, Inclusion AI
Y
Yuxin Ma
AGI Research Center, Inclusion AI
Shuheng Shen
Shuheng Shen
Ant Group
Machine LearningOptimizationPrivacy
P
Peng Sun
AGI Research Center, Inclusion AI
C
Chaoyang Wang
AGI Research Center, Inclusion AI
H
Hongjun Wang
AGI Research Center, Inclusion AI