Benchmarking General Mobile Assistants in Challenging Real-World Scenarios

📅 2026-08-21
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文通过引入GMA基准,涵盖更多应用和复杂任务,评估移动助手在真实场景中的表现,并研究了不同设计对性能的影响。
📝 Abstract
Graphical user interfaces have emerged as an important environment for evaluating autonomous AI agents on multimodal interactive tasks. Existing benchmarks such as AndroidWorld and MobileWorld provide strong foundations for mobile agent evaluation, but their application coverage and task design do not yet fully capture the diversity and complexity of realistic mobile use. We present GMA, a benchmark for evaluating general mobile assistants in challenging real-world scenarios. GMA introduces seven applications based on open-source projects, spanning domains such as lifestyle sharing and travel planning, and 300 tasks across four difficulty tiers, from atomic actions to complex multi-step workflows. We evaluate eight frontier models and find that performance declines substantially as task complexity increases, with current agents remaining far from reliably handling realistic user requirements. We further conduct controlled ablation studies of agent harness choices, including context retention and explicit state tracking, under a shared environment, model setting, and task taxonomy. Results show that appropriate harness design can meaningfully improve performance, particularly on demanding workflows, while the effectiveness of specific designs can vary across foundation models. Overall, GMA complements existing benchmarks by expanding application coverage and task complexity, providing a challenging testbed for evaluating mobile agents and studying how harness design supports reliable execution in complex mobile workflows.
Problem

Research questions and friction points this paper is trying to address.

Benchmarking
General Mobile Assistants
Real-World Scenarios
Task Complexity
Application Coverage
Innovation

Methods, ideas, or system contributions that make the work stand out.

General Mobile Assistants
Real-World Scenarios
Task Complexity
Harness Design
Performance Improvement
🔎 Similar Papers
Yiqi Zhu
Yiqi Zhu
Undergraduate Student, Tsinghua University
Artificial Intelligence
F
Feiyu Gao
Alibaba Token Hub, Alibaba Group
J
Jiaxing Fan
Alibaba Token Hub, Alibaba Group
J
Jiahui Zeng
Alibaba Token Hub, Alibaba Group
M
Minggang Wu
Alibaba Token Hub, Alibaba Group
Chenliang Li
Chenliang Li
Alibaba Inc.
nlp
Haiyang Xu
Haiyang Xu
Alibaba Group, DIDI AI LABS, SEU
Multimodal LearningLarge Language ModelAgentNatural Language Processing
P
Peng Li
Institute for AI Industry Research (AIR), Tsinghua University, Beijing, China
Ming Yan
Ming Yan
Alibaba Group
Y
Yang Liu
Dept. of Comp. Sci. & Tech., Institute for AI, Tsinghua University, Beijing, China; Institute for AI Industry Research (AIR), Tsinghua University, Beijing, China