WM-R1: Training GUI Agents to Reason and leverage World Models with Reinforcement Learning

📅 2026-08-27
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决GUI代理训练中高资源成本与不稳定性问题,提出WM-R1框架,利用世界模型代替真实环境进行强化学习,提高训练效率和性能。
📝 Abstract
GUI agents trained with reinforcement learning (RL) have showcased strong environment learning capabilities on mobile platforms. However, RL typically demands extensive real-environment interactions, leading to high resource costs and instability, especially in GUI scenarios. To address these, we propose WM-R1, the first reinforcement learning framework that trains mobile GUI agents with world models instead of real environments. Specifically, world models serve as the source of state transitions during all rollouts, replacing the real Android environment within the training loop. WM-R1 also embeds world models directly into the thinking process, enabling agents to reason about the consequences of candidate actions before committing to the final action. Crucially, WM-R1 eliminates the need for real-environment interaction, supports massively parallelized and step-level granularized trajectory generation grounded in world models, and introduces a multi-dimensional rule-based reward that jointly optimizes task success, trajectory efficiency, and world model utilization. For efficient training, we curate a high-quality dataset of 2000 challenging tasks. Experiments on Android mobile benchmarks demonstrate that WM-R1-trained agents significantly outperform GRPO-only baselines and inference-time simulation methods. Code is available at https://github.com/genalyu/WM-R1 .
Problem

Research questions and friction points this paper is trying to address.

reinforcement learning
GUI agents
real-environment interactions
resource costs
instability
Innovation

Methods, ideas, or system contributions that make the work stand out.

World Models
Reinforcement Learning
Mobile GUI Agents
Parallelized Trajectory Generation
Rule-based Reward
💼 Related Jobs
No related jobs found.
Y
Yu Han
School of Computer Science and Technology, East China Normal University
Tianwen Qian
Tianwen Qian
East China Normal University
MultimediaVision and LanguageEmbodied AI