BlueLM-GUI Technical Report: A Real-Device-Centric Flywheel for Self-Improving Mobile GUI Agents

📅 2026-09-10
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出BlueLM-GUI,通过三个原则解决移动GUI代理在工业部署中的分布不匹配、真机利用不足和基准饱和问题,提升模型能力。
📝 Abstract
Mobile GUI agents are shifting from multi-module frameworks to native models trained end-to-end, yet industrial deployment faces three persistent gaps. Sandbox training produces a distribution mismatch with production environments; expensive real-device failures remain underutilized; and fixed benchmarks saturate, losing the power to guide iteration. We present BlueLM-GUI, a 35B-A3B mobile GUI agent built as a real-device-centric flywheel that closes these gaps through three principles. Every Sample Matters: a dual-track pipeline with Heterogeneous Triple-System Consensus evaluation and an Error Correction \& Derivation Module salvages every trajectory into usable supervision. Every Rollout Is Real: a three-stage recipe---continual pre-training, supervised fine-tuning, and agentic reinforcement learning on hundreds of real phones---grounds every rollout in real production environments, so the capability the model learns transfers directly to deployment. Every Query Evolves: a quota-driven benchmark methodology with three orthogonal axes enables precise attribution and allows the benchmark to be systematically upgraded as the model improves. BlueLM-GUI achieves 87.4 on MobileGUI-VBench, surpassing the best closed-source model by 5.1 points, and 84.9 on AndroidWorld, the best result among open-source models and competitive with closed-source models. These results demonstrate that grounding model training and iterative improvement in both real devices and the three Every principles yields strong, robust, and transferable mobile GUI capability.
Problem

Research questions and friction points this paper is trying to address.

mobile GUI agents
distribution mismatch
real-device failures
fixed benchmarks
Innovation

Methods, ideas, or system contributions that make the work stand out.

Real-Device-Centric
Heterogeneous Triple-System Consensus
Error Correction & Derivation Module
Quota-Driven Benchmark Methodology
T
Tong Ye
vivo AI Lab
K
Kunyang Han
vivo AI Lab
G
Guozhi Wang
vivo AI Lab
L
Longqiang Luo
vivo AI Lab
Z
Zhifeng Ding
vivo AI Lab
Y
Yongxiang Zhang
vivo AI Lab
X
Xiaolei Shen
vivo AI Lab
Y
Yuxuan Zhang
vivo AI Lab
Z
Zhuping Zhang
vivo AI Lab
T
Tao Xu
vivo AI Lab
Y
Yue Pan
vivo AI Lab
Yucheng Zhao
Yucheng Zhao
MEGVII Technology
RobotLarge Language ModelVideo Generation
Y
Yupei Hu
vivo AI Lab
Y
Yuanjiang Ouyang
vivo AI Lab
D
Danfeng Shen
vivo AI Lab
Runqi Lin
Runqi Lin
PhD student, The University of Sydney
Machine LearningAI SafetyTrustworthy MLAdversarial Robustness
H
Hongda Cai
vivo AI Lab
Z
Zhaoxiong Wang
vivo AI Lab
Mengjia Yan
Mengjia Yan
vivo AI Lab
Y
Yingjie Zhong
vivo AI Lab
C
Chen Zhou
vivo AI Lab
Z
Zeyu Zhang
vivo AI Lab
X
Xuwen Zhu
vivo AI Lab
P
Penggang Shi
vivo AI Lab
M
Mingcheng Luo
vivo AI Lab