Swift-Image: Exploring the Performance Frontier of Compact Unified Image Generation Models

📅 2026-08-20
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过系统训练工程和优化方法,探讨了在有限计算资源下小型图像生成模型的性能极限,提出了Swift-Image模型。
📝 Abstract
We present Swift-Image, a compact unified model for text-to-image generation, single-image editing, and multi-image editing. Our goal is to explore how far a relatively small visual generator can be pushed through systematic training engineering under a constrained computational budget. Swift-Image adopts an efficient 6B single-stream DiT and a progressive training pipeline that evolves from broad semantic coverage to higher resolution, stronger visual quality, and unified generation-editing supervision. For post-training, we employ parallel expert reinforcement learning followed by multi-teacher on-policy distillation to alleviate interference among heterogeneous objectives. We further decouple high-level reasoning from pixel-level rendering with a Prompt Enhancer that translates user requests into generator-aligned visual specifications. For efficient deployment, structural pruning and few-step distillation produce 3B and accelerated variants. Swift-Image achieves leading aggregate performance among evaluated open-source models with only 6B parameters and 243K GPU training hours; the compressed 3B model incurs nearly no loss, while few-step distillation further improves aggregate editing performance with substantially fewer sampling steps. Our study also summarizes practical lessons for architecture, data curriculum, post-training, prompt enhancement, and model compression.
Problem

Research questions and friction points this paper is trying to address.

Compact Unified Model
Text-to-Image Generation
Single-Image Editing
Multi-Image Editing
Constrained Computational Budget
Innovation

Methods, ideas, or system contributions that make the work stand out.

compact unified model
progressive training pipeline
parallel expert reinforcement learning
multi-teacher on-policy distillation
prompt enhancer
🔎 Similar Papers
No similar papers found.
Taihang Hu
Taihang Hu
Nankai University
Deep LearningComputer VisionGenerative models
Z
Zhao Wang
Alibaba Group
Zuan Gao
Zuan Gao
University of Science and Technology
GenAIAIGCOCR
T
Tao Liu
Alibaba Group
H
Hao Yan
Alibaba Group
Z
Zhengze Xu
Alibaba Group
Y
Yuhang Yu
Alibaba Group
Y
Yongchao Du
Alibaba Group
X
Xingjian Wang
Alibaba Group
J
Jun Zheng
Alibaba Group
Q
Qinye Zhou
Alibaba Group
Z
Zhengrui Chen
Alibaba Group
C
Chao Lin
Alibaba Group
Y
Yefeng Shen
Alibaba Group
Z
Zhengtao Wu
Alibaba Group
G
Ge Wu
Alibaba Group
Xiaoli Xu
Xiaoli Xu
Southeast University, China
Wireless communicationnetwork codingchannel coding
D
Denghui Yang
Alibaba Group
Huayu Zhang
Huayu Zhang
Senior Engineer, Huawei Technologies Co., Ltd
Distributed SystemNetwork ScienceMachine LearningOptimizationGraph Theory
M
Mingzhou Zhang
Alibaba Group
Mengting Chen
Mengting Chen
Alibaba Group
Generative ModelingComputer Vision