WebChallenger: A Reliable and Efficient Generalist Web Agent

๐Ÿ“… 2026-06-09
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This work addresses the high reasoning cost, poor reusability, and lack of human-like cognitive mechanisms in general-purpose large language model agents for autonomous web navigation, particularly their inefficiency in repetitive tasks. The authors propose a model-size-agnostic agent framework that emulates human cognitive strengths through architectural design: it introduces PageMem, a DOM-based structured page representation, and integrates three key mechanismsโ€”divide-and-conquer observation (mimicking selective attention), a lightweight exploration memory system (enabling structured memory), and a composite action workflow (enhancing operational fluency). Without fine-tuning any open-source large language models, the approach achieves success rates of 56.3%, 48.7%, 51.0%, and 70.9% on WebArena, VisualWebArena, Online-Mind2Web, and WorkArena, respectively, matching the performance of state-of-the-art closed-source systems while significantly reducing computational costs.
๐Ÿ“ Abstract
Autonomous web navigation remains challenging for LLM agents, and the strongest generalist systems rely on proprietary reasoning models whose inference cost is prohibitive for the repetitive tasks where such agents would be most useful. We argue this gap stems not from insufficient model capability but from agent architectures that fail to replicate three human cognitive advantages: selective attention to relevant page regions, persistent memory of website structure, and procedural fluency with common interaction patterns. We introduce WebChallenger, a web agent framework that addresses each gap through architecture design rather than model scale, built around PageMem: a structured page representation deterministically constructed from the DOM that exposes each page as a hierarchy of semantic sections with short summaries. On this shared substrate we build three mechanisms that mirror the three cognitive advantages: a divide-and-conquer observation pipeline that lets the agent skim section summaries and extract details only from task-relevant regions; a lightweight exploration and memory system that traverses each website once to build a reusable map of pages and element behaviors; and compound action workflows that collapse common multi-step interactions into single agent actions, handling partial state changes automatically. Because all three operate over PageMem, the framework generalizes across websites without site-specific adapters. Using off-the-shelf open-weight models without fine-tuning, our system achieves 56.3% on WebArena, 48.7% on VisualWebArena, 51.0% on Online-Mind2Web, and 70.9% on WorkArena, approaching frontier proprietary systems at a fraction of the cost. Our code is released at https://github.com/jayoohwang1/webchallenger
Problem

Research questions and friction points this paper is trying to address.

autonomous web navigation
LLM agents
cognitive advantages
web agent architecture
inference cost
Innovation

Methods, ideas, or system contributions that make the work stand out.

web agent
structured page representation
cognitive-inspired architecture
procedural fluency
cross-site generalization
๐Ÿ”Ž Similar Papers
No similar papers found.