Orcust: Stepwise-Feedback Reinforcement Learning for GUI Agent
Existing GUI agents face two critical challenges: unreliable reward signals and limited capacity for online trajectory generation—leading to unreliable reasoning and low data efficiency. To address these, we propose Orcust, a novel framework introducing two key mechanisms: (1) Principle-Constrained Reward Modeling (PCRM), which incorporates interpretable, human-aligned rules into the reinforcement learning reward function to ensure policy adherence to domain principles; and (2) Online Virtual Machine-Grounded Trajectory Construction (OVTC), which leverages a lightweight virtual machine to autonomously generate high-quality, structured interaction trajectories enabling fine-grained, stepwise training. Integrating chain-of-thought reasoning with LLM-driven rule feedback, Orcust achieves +22.2% and +23.9% improvements on ScreenSpot and ScreenSpot-Pro benchmarks, respectively. The framework significantly enhances reasoning reliability, task adaptability, and system scalability.