GUI-Eyes: Tool-Augmented Perception for Visual Grounding in GUI Agents

📅 2026-01-14
📈 Citations: 1
Influential: 0
📄 PDF
🤖 AI Summary
This work proposes GUI-Eyes, a reinforcement learning–based active visual perception framework that addresses the limitations of existing GUI agents, which predominantly rely on static visual inputs and lack the ability to actively decide when and how to observe the interface. GUI-Eyes employs a two-stage policy network to autonomously invoke visual tools—such as cropping and zooming—to enable progressive perception from coarse exploration to fine-grained localization. The approach innovatively integrates a vision-language model with a tool-calling mechanism and introduces a continuous-space reward function that combines positional proximity and region overlap, effectively mitigating reward sparsity in GUI environments. Evaluated on the ScreenSpot-Pro benchmark, GUI-Eyes-3B achieves a localization accuracy of 44.8% using only 3k annotated samples, significantly outperforming current supervised and reinforcement learning baselines.

Technology Category

Application Category

📝 Abstract
Recent advances in vision-language models (VLMs) and reinforcement learning (RL) have driven progress in GUI automation. However, most existing methods rely on static, one-shot visual inputs and passive perception, lacking the ability to adaptively determine when, whether, and how to observe the interface. We present GUI-Eyes, a reinforcement learning framework for active visual perception in GUI tasks. To acquire more informative observations, the agent learns to make strategic decisions on both whether and how to invoke visual tools, such as cropping or zooming, within a two-stage reasoning process. To support this behavior, we introduce a progressive perception strategy that decomposes decision-making into coarse exploration and fine-grained grounding, coordinated by a two-level policy. In addition, we design a spatially continuous reward function tailored to tool usage, which integrates both location proximity and region overlap to provide dense supervision and alleviate the reward sparsity common in GUI environments. On the ScreenSpot-Pro benchmark, GUI-Eyes-3B achieves 44.8% grounding accuracy using only 3k labeled samples, significantly outperforming both supervised and RL-based baselines. These results highlight that tool-aware active perception, enabled by staged policy reasoning and fine-grained reward feedback, is critical for building robust and data-efficient GUI agents.
Problem

Research questions and friction points this paper is trying to address.

GUI automation
active visual perception
visual grounding
tool-augmented perception
reinforcement learning
Innovation

Methods, ideas, or system contributions that make the work stand out.

active perception
tool-augmented vision
two-stage reasoning
spatially continuous reward
GUI grounding
Chen Chen
Chen Chen
Postdoc@UChicago, PhD@UIUC, BS@USTC
Polymeric MaterialsPolymer PhysicsPolymer Electrolytes
J
Jiawei Shao
Institute of Artificial Intelligence (TeleAI), China Telecom
D
Dakuan Lu
Institute of Artificial Intelligence (TeleAI), China Telecom
H
Haoyi Hu
Shanghai Jiao Tong University
X
Xiangcheng Liu
University of Science and Technology of China
H
Hantao Yao
University of Science and Technology of China
W
Wu Liu
University of Science and Technology of China