Act2Intention: A Benchmark For Developing Active Mobile Agents Through Inferring User Intention from GUI Actions

📅 2026-08-14
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the lack of proactive service capabilities in mobile GUI agents by proposing a three-stage framework that integrates understanding, prediction, and execution. We introduce the first continuous intent-action trajectory benchmark, comprising 70,000 intents and 700,000 actions, to facilitate rigorous evaluation. Through supervised fine-tuning of multimodal large language models and personalized prediction algorithms, this work achieves a paradigm shift from passive response to proactive service. Experimental results demonstrate that the proposed framework improves accuracy in intent understanding, prediction, and execution by 32.0%, 10.25%, and 6.9%, respectively. These findings indicate a significant enhancement in the proactive interaction performance of mobile GUI agents, establishing a new baseline for autonomous mobile assistance.
📝 Abstract
Mobile GUI Agents powered by multimodal large language models (MLLMs) show promise in human-computer intelligence. However, current research primarily focuses on reactive task execution while lacking a comprehensive understanding-prediction-execution process for user intentions, which are the core requirements of active agents. In this paper, we propose the Act2Intention framework that builds an active mobile agent by integrating understanding, predicting user intentions, and executing decisions. First, we construct the Act2Intention Bench through data collection and validated generation, comprising 72,511 intentions and over 700,000 actions across 52 apps, thereby establishing the first benchmark for evaluating proactive agents via continuous intention-action trajectories. We further develop the Act2Intention Agent, achieving proactive services through Proactive-oriented Intention Understanding, Personalized Proactive Intention Prediction, and Experience-guided Intention Execution. Experimental results show that supervised fine-tuning on Act2Intention Bench yields absolute improvements of +32.0 Acc-S, +10.25 Acc-S, and +6.9 SSR points over non-fine-tuned counterparts under the same agent framework for intention understanding, prediction, and execution, respectively. This success underscores the necessity and value of the Act2Intention Bench, which establishes a standardized platform for developing and evaluating proactive agents and consequently paves the way for research on intention-driven human-computer interaction.
Problem

Research questions and friction points this paper is trying to address.

Active Mobile Agents
User Intention Inference
Proactive Agents
Benchmark
Mobile GUI
Innovation

Methods, ideas, or system contributions that make the work stand out.

Active Mobile Agents
User Intention Inference
Act2Intention Benchmark
Proactive Intention Prediction
GUI Actions
🔎 Similar Papers
2024-06-20arXiv.orgCitations: 3