Institution profile

Simular

Industry research
Official website
Research library2linked papers
Opportunities0open roles
Selected work

Representative Papers

The Unreasonable Effectiveness of Scaling Agents for Computer Use

Oct 02, 2025

Current computer-using agents (CUAs) exhibit low reliability and high performance variance on long-horizon, complex digital tasks. To address this, we propose Behavior Best-of-N (bBoN), the first framework to integrate scalable agent architectures with behavioral narrative modeling: it generates diverse execution trajectories via multi-agent rollouts, employs behavior narratives for structured trajectory modeling and evaluation, and introduces a reinforcement learning–driven selection mechanism. bBoN significantly improves robustness and cross-platform generalization, achieving 69.9% task success rate on OSWorld—approaching human performance (72%)—and is validated on WindowsAgentArena and AndroidWorld, establishing new state-of-the-art results. Its core contribution lies in establishing a behavior-centric paradigm for scalable CUAs, providing an extensible technical pathway toward reliable, general-purpose computer-use automation.

0 citationsRead paper

Agent S2: A Compositional Generalist-Specialist Framework for Computer Use Agents

Apr 01, 2025

To address three key challenges in GUI agents—imprecise element grounding, difficulty in long-horizon task planning, and limited cognitive capabilities of general-purpose models—this paper proposes a novel computer-use agent specifically designed for GUI interaction. Methodologically: (1) We introduce Mixture-of-Grounding, a hybrid localization technique that enhances fine-grained perception of GUI elements; (2) we design Proactive Hierarchical Planning to enable multi-scale, dynamic task decomposition and reconstruction; and (3) we establish a collaborative architecture integrating general-purpose foundation models with specialized modules for visual grounding, hierarchical reasoning, real-time observational feedback, and cross-platform adaptation. Evaluated on OSWorld (15/50-step), WindowsAgentArena, and AndroidWorld, our agent achieves comprehensive performance gains over state-of-the-art baselines—improving success rates by 18.9%–52.8%—and demonstrates significantly enhanced cross-OS generalization capability.

0 citationsRead paper
Recent publications

Latest Papers

The Unreasonable Effectiveness of Scaling Agents for Computer Use

Oct 02, 2025

Current computer-using agents (CUAs) exhibit low reliability and high performance variance on long-horizon, complex digital tasks. To address this, we propose Behavior Best-of-N (bBoN), the first framework to integrate scalable agent architectures with behavioral narrative modeling: it generates diverse execution trajectories via multi-agent rollouts, employs behavior narratives for structured trajectory modeling and evaluation, and introduces a reinforcement learning–driven selection mechanism. bBoN significantly improves robustness and cross-platform generalization, achieving 69.9% task success rate on OSWorld—approaching human performance (72%)—and is validated on WindowsAgentArena and AndroidWorld, establishing new state-of-the-art results. Its core contribution lies in establishing a behavior-centric paradigm for scalable CUAs, providing an extensible technical pathway toward reliable, general-purpose computer-use automation.

0 citationsRead paper

Agent S2: A Compositional Generalist-Specialist Framework for Computer Use Agents

Apr 01, 2025

To address three key challenges in GUI agents—imprecise element grounding, difficulty in long-horizon task planning, and limited cognitive capabilities of general-purpose models—this paper proposes a novel computer-use agent specifically designed for GUI interaction. Methodologically: (1) We introduce Mixture-of-Grounding, a hybrid localization technique that enhances fine-grained perception of GUI elements; (2) we design Proactive Hierarchical Planning to enable multi-scale, dynamic task decomposition and reconstruction; and (3) we establish a collaborative architecture integrating general-purpose foundation models with specialized modules for visual grounding, hierarchical reasoning, real-time observational feedback, and cross-platform adaptation. Evaluated on OSWorld (15/50-step), WindowsAgentArena, and AndroidWorld, our agent achieves comprehensive performance gains over state-of-the-art baselines—improving success rates by 18.9%–52.8%—and demonstrates significantly enhanced cross-OS generalization capability.

0 citationsRead paper