LLaDA-UI: Bringing Block-wise Diffusion to Vision-Language GUI Agents

📅 2026-09-09
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究解决了如何将扩散大语言模型应用于GUI代理的问题,通过开发LLaDA-UI,采用两阶段训练方法,实现了高效且准确的多模态GUI交互。
📝 Abstract
Diffusion large language models (dLLMs) achieve high decoding efficiency through block-parallel, arbitrary-order generation, making them attractive for latency-sensitive applications. GUI agents represent a natural testbed for this paradigm, as they must repeatedly perceive screen states and emit structured, spatially grounded actions in real time. However, whether dLLMs can be extended into capable multimodal GUI agents while preserving their parallel decoding advantage remains an open question. We present LLaDA-UI, a 16.7B-parameter MoE-based, block-wise diffusion vision-language GUI agent. LLaDA-UI follows a two-stage training pipeline: general multimodal pre-training aligns a native-resolution vision encoder with the LLaDA2.0-mini-base diffusion language backbone, followed by GUI-agent supervised fine-tuning on diverse mobile, desktop, web, and grounding data. Across widely adopted grounding benchmarks and navigation benchmarks spanning multiple platforms, LLaDA-UI substantially outperforms Qwen2.5-VL-7B and surpasses Qwen3-VL-8B on four of six reported GUI benchmarks. These results establish block-wise diffusion as a practical generative paradigm for multimodal GUI agents.
Problem

Research questions and friction points this paper is trying to address.

diffusion large language models
multimodal GUI agents
block-parallel generation
real-time perception
structured actions
Innovation

Methods, ideas, or system contributions that make the work stand out.

block-wise diffusion
multimodal GUI agents
MoE-based
Zhangxuan Gu
Zhangxuan Gu
Ant Group
computer vision
H
Haoxing Chen
AGI Research Center, Inclusion AI
Q
Qi Qin
AGI Research Center, Inclusion AI
Yi Xin
Yi Xin
California Institute of Technology
Industrial OrganizationEconometrics
K
Kai Gan
AGI Research Center, Inclusion AI
L
Lin Liu
AGI Research Center, Inclusion AI
L
Long Cui
AGI Research Center, Inclusion AI
X
Xiaomei Wang
AGI Research Center, Inclusion AI
Beitong Zhou
Beitong Zhou
Huazhou University of Science and Technology
deep learningcomputer vision
Y
Yunzhu Zhang
Venus Team
Z
Zhengwen Zeng
Venus Team
C
Changlong Gao
Venus Team
W
Weizhi Chen
Venus Team
R
Rongchao Zhang
Venus Team
Haoyuan Wu
Haoyuan Wu
The Chinese University of Hong Kong
Generative AILarge Language ModelsMultimodal ModelsAgentic AIRepresentation Learning
Shuheng Shen
Shuheng Shen
Ant Group
Machine LearningOptimizationPrivacy
C
Changhua Meng
Venus Team
W
Weiqiang Wang
Venus Team
Jianguo Li
Jianguo Li
Director, Ant Group
deep learningcomputer visionmachine learningsystem
Zhenzhong Lan
Zhenzhong Lan
School of Engineering, Westlake University
NLPComputer VisionMultimedia