UI-Venus-2 Technical Report

📅 2026-08-27
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出UI-Venus-2,通过扩展环境覆盖、改进任务构建和增强验证机制,解决多模态GUI代理在实际应用中的可靠性问题。
📝 Abstract
Multimodal GUI agents have emerged as a promising paradigm for digital task automation, yet transitioning from benchmark-oriented models to dependable real-world applications remains challenging due to limited environment coverage, brittle task construction, and unreliable reward verification. In this work, we present UI-Venus-2, a general-purpose foundation GUI agent designed to operate across mobile, web, and desktop environments through a unified closed-loop reasoning-action framework. To bridge the gap toward practical deployment, we jointly scale three critical dimensions: (1) Environments, expanding coverage to more than 170 multilingual mobile apps and native desktop operating systems; (2) Tasks, employing a deep-research pipeline for function-grounded instruction generation; and (3) Verification, adopting trace-level and sample-level evaluators with visual keypoints and multi-model voting to ensure reliable RL signals for training. Furthermore, we integrate safety-aware mechanisms to ensure controlled execution of consequential actions. By offering a capable, efficient, and open-source foundation, UI-Venus-2 advances the field toward more generalizable, verifiable, and self-reflective agents for real-world applications.
Problem

Research questions and friction points this paper is trying to address.

Multimodal GUI agents
digital task automation
environment coverage
brittle task construction
unreliable reward verification
Innovation

Methods, ideas, or system contributions that make the work stand out.

multimodal GUI agents
unified closed-loop reasoning-action framework
function-grounded instruction generation
multi-model voting
🔎 Similar Papers
💼 Related Jobs
No related jobs found.