UI-CUBE: Enterprise-Grade Computer Use Agent Benchmarking Beyond Task Accuracy to Operational Reliability
Existing CUA (Computer-User-Agent) benchmarks overemphasize functional correctness and fail to assess agent reliability in enterprise production environments. Method: We propose UI-CUBE—a first-of-its-kind, enterprise-ready, systematic diagnostic benchmark comprising 226 graded tasks. It integrates UI perturbation, multi-resolution UI testing, and application-state verification to rigorously evaluate architectural deficiencies in memory management, hierarchical planning, and state coordination. Contribution/Results: Experiments reveal that state-of-the-art agents achieve only 67–85% success on simple tasks but plummet to 9–19% on complex enterprise workflows—far below human novices (61.2%), exposing fundamental architectural bottlenecks. UI-CUBE establishes a reproducible, attributable reliability evaluation paradigm for industrial deployment of CUAs.