Efficient GUI Agents: A Systems Survey of Observation, Memory, Action, and Runtime Optimization

๐Ÿ“… 2026-09-02
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
ๆœฌๆ–‡้€š่ฟ‡็ณป็ปŸ่ง’ๅบฆ็ ”็ฉถไบ†GUIไปฃ็†็š„ๆ•ˆ็އ้—ฎ้ข˜๏ผŒๅŒ…ๆ‹ฌ่ง‚ๅฏŸใ€่ฎฐๅฟ†ใ€่กŒๅŠจๅ’Œ่ฟ่กŒๆ—ถไผ˜ๅŒ–๏ผŒๅนถๆ€ป็ป“ไบ†ๆ้ซ˜ๆ•ˆ็އ็š„ไธป่ฆๆ–นๆณ•ใ€‚
๐Ÿ“ Abstract
GUI agents increasingly operate across websites, mobile apps, and desktop environments, yet the field still reports progress primarily through task success. We argue that practical deployment depends equally on efficiency: how much context, computation, action budget, and runtime overhead an agent consumes while succeeding. This survey studies efficient GUI agents through an end-to-end systems lens that preserves the current technical axes of observation efficiency, context and memory efficiency, action efficiency, and planner-side/system efficiency. For each subsection, we expand the seed literature through targeted search plus backward and forward citation chaining, then synthesize the dominant mechanisms, reported efficiency signals, and new overheads they introduce. Across the literature, recent progress converges on a small set of recurring ideas: selective reading instead of full-context ingestion, global-to-local visual allocation, recoverable memory rather than raw history replay, verification-aware control, and hybrid runtimes that can switch between GUI and non-GUI execution. We conclude by identifying the main open problems, including honest accounting of verifier cost, cross-benchmark comparability, and co-design of observation, memory, and execution layers under real latency and privacy constraints.
Problem

Research questions and friction points this paper is trying to address.

GUI agents
efficiency
context
computation
runtime overhead
Innovation

Methods, ideas, or system contributions that make the work stand out.

selective reading
global-to-local visual allocation
recoverable memory
verification-aware control
hybrid runtimes
๐Ÿ”Ž Similar Papers
No similar papers found.
B
Bizhe Bai
College of Future Information Technology, Fudan University, Shanghai, China
Jiakang Yuan
Jiakang Yuan
Fudan university
MLLMsMulti-agent SystemReasoning
H
Hongming Wu
College of Future Information Technology, Fudan University, Shanghai, China
X
Xinyue Wang
College of Future Information Technology, Fudan University, Shanghai, China
J
Jie Ren
College of Future Information Technology, Fudan University, Shanghai, China
S
Siyao Chen
College of Future Information Technology, Fudan University, Shanghai, China
Y
Yuchen Ya
College of Future Information Technology, Fudan University, Shanghai, China
F
Fan Bai
Independent Researcher
P
Pai Peng
Independent Researcher
Huafeng Qin
Huafeng Qin
Chongqing Technology and Business University
Biometrics (e.g veinface and gait)computer visionand machine learning
Tao Chen
Tao Chen
Fudan University
Deep LearningMedical Image Segmentation