Mano Report
GUI automation remains challenged by visual complexity, environmental dynamics, and multi-step reasoning. Existing vision-language model (VLM) approaches suffer from low-resolution input, domain shift, and weak sequential decision-making capabilities. This paper proposes a multimodal foundation model–based GUI agent featuring a three-stage collaborative training pipeline—pretraining, domain-specific fine-tuning, and reinforcement-based refinement—alongside a validation-driven error recovery mechanism. We further construct a high-fidelity simulation environment to generate high-quality interactive data. Our approach innovatively integrates multimodal pretraining, offline and online reinforcement learning, cross-domain transfer, and interpretable action modeling. On the Mind2Web and OSWorld benchmarks, our method achieves task success rates of 82.4% and 76.9%, respectively—significantly surpassing state-of-the-art methods—and marks the first unified framework for robust, recoverable, and multi-step GUI automation.