MobileExplorer: Accelerating On-Device Inference for Mobile GUI Agents via Online Exploration

📅 2026-05-26

📈 Citations: 0

✨ Influential: 0

career value

194K/year

🤖 AI Summary

This work addresses the limitations of existing mobile GUI agents, which rely on cloud-based inference and suffer from privacy risks and high latency, hindering efficient on-device deployment. The authors propose the first on-device GUI agent framework capable of online exploration, featuring concurrent lightweight UI probing during vision-language model inference to construct structured memory. A dual-layer rollback mechanism is introduced to ensure operational reliability, while exploration trajectories are distilled into contextual prompts to accelerate subsequent reasoning. Experiments on AndroidWorld and complex dynamic tasks demonstrate that the proposed approach reduces end-to-end latency by 23%, improves task success rates by up to 5%, and significantly decreases the average number of reasoning steps.

📝 Abstract

Mobile graphical user interface (GUI) agents enable AI models to autonomously operate smartphones on behalf of users. However, most existing systems focus primarily on optimizing task accuracy and rely on cloud-hosted models for inference, which introduces privacy concerns and network-dependent latency. As a result, fully on-device deployment of mobile GUI agents remains underexplored. We propose MobileExplorer, a new framework that accelerates on-device inference for vision-based mobile GUI agents via online exploration. The key idea is to exploit the long per-step reasoning time of vision-language models (VLMs) by performing lightweight, parallel exploration of UI elements. During model inference, the agent proactively probes semantically relevant UI elements and records these exploration traces as structured memory. To ensure reliable execution in live mobile environments, we design a two-level rollback mechanism that robustly restores the initial UI state when a fast but naive backtracking strategy fails. The collected exploration traces are then summarized into concise contextual hints and injected into the prompt to enhance the subsequent reasoning step. We evaluate MobileExplorer on multiple off-the-shelf devices using the AndroidWorld benchmark, as well as newly designed, more complex tasks and dynamic on-device environments. MobileExplorer reduces the average number of reasoning steps and end-to-end latency by 23\%, while maintaining or improving task success rates by up to 5\%. A video demonstration of MobileExplorer performance in the real world is available at https://youtu.be/thK7MJmdlvM .

Problem

Research questions and friction points this paper is trying to address.

on-device inference

mobile GUI agents

privacy concerns

network latency

vision-language models

Innovation

Methods, ideas, or system contributions that make the work stand out.

on-device inference

online exploration

vision-language models