🤖 AI Summary
This study addresses the pervasive challenges in real-world deployment of GUI agent systems—namely technical fragility, weak engineering foundations, and insufficient validation—stemming from a lack of comprehensive software engineering support across their lifecycle. Synthesizing insights from 336 studies published between 2018 and 2026, this work introduces a software engineering perspective to the field for the first time, employing bibliometric and qualitative analyses centered on the perception–reasoning–action closed-loop architecture. It critically examines key elements including modular design, interactive evaluation, recovery mechanisms, and human-agent collaboration. The analysis reveals architectural imbalances and evaluation limitations, noting that despite rapid growth since 2024, engineering practices have lagged. To bridge this gap, the paper proposes an engineering framework integrating observability, privacy preservation, and human-centered governance, emphasizing auditability, security controls, and post-deployment maintenance to realize reliable, maintainable, and deployable GUI agent systems.
📝 Abstract
GUI agents have advanced rapidly, producing a growing body of frameworks, benchmarks, and applications. However, this growth has outpaced the maturity of the field. GUI agents remain technically brittle, incompletely engineered, and insufficiently validated for sustained real-world use. They are evolving into closed-loop software systems. Within these systems, model reasoning is coupled with interface perception, execution feedback, recovery, and human oversight. This evolution calls for a software engineering perspective that remains largely absent from existing research. We address this gap by reviewing 336 GUI-agent papers from January 2018 to April 2026. Five research questions examine the research landscape, architectures, evaluation, software lifecycle concerns, and future opportunities. Our findings show that the field has expanded sharply since 2024, while mobile and web settings remain dominant. Architectures increasingly adopt modular perceive-reason-act loops, but recovery, human escalation, safety enforcement, and auditability remain underdeveloped. This architectural imbalance extends to evaluation. Evaluations are becoming more interactive, but they remain centered on task success and are difficult to compare across protocols. More broadly, existing studies provide limited support for testing beyond benchmarks and for maintaining agents after release. Observability, privacy engineering, and systematic human oversight are also underdeveloped. Together, these findings show that capability improvements alone cannot ensure deployment readiness. Future research should connect dependable execution with lifecycle-centered testing and reproducible evaluation. It should also integrate permission and privacy controls with cost-aware, human-centered governance. This integration is necessary to build dependable, maintainable, secure, and deployable GUI-agent systems.