Institution profile

Hcompany

Research institution
Research library5linked papers
Opportunities0open roles
Selected work

Representative Papers

DragOn: A Benchmark and Dataset for Drag-Based GUI Interactions

Jun 04, 2026

This work addresses the limited performance of existing GUI agents in drag-based interactions—such as text highlighting and slider manipulation—which stems primarily from the absence of large-scale, high-quality datasets capturing such operations. To bridge this gap, the authors present DragOn, the first systematically constructed benchmark dataset specifically designed for drag interactions, encompassing four representative scenarios with 286,000 training screenshots, 3.5 million task instances, and 2,000 evaluation samples. They further introduce an end-to-end drag localization training paradigm conditioned on screen images and natural language instructions. Fine-tuning the Qwen vision-language model on DragOn substantially improves performance on drag-related tasks, demonstrating the dataset’s effectiveness in enhancing model generalization to real-world digital interactions and filling a critical data void in complex, continuous GUI manipulation.

0 citationsRead paper

Best-of-Q: Improving VLM agents with Q-function Action Ranking at Inference

Jan 30, 2026

Existing vision-language model (VLM) agents exhibit limited adaptability in dynamic environments such as web navigation and incur high costs when fine-tuned. This work proposes a training-free, inference-time optimization approach that decouples action generation from selection: the VLM is frozen to generate candidate actions, while a lightweight, offline-trained Q-function reranks these candidates. Notably, this is the first method to directly employ a Q-function for action selection during inference without updating the underlying policy, enabling immediate performance gains. Evaluated on the WebVoyager benchmark, the approach boosts the success rate of Qwen2.5-VL-7B from 38.8% to 55.7% and GPT-4.1 from 82.4% to 88.8%, significantly outperforming baseline methods.

0 citationsRead paper

Surfer 2: The Next Generation of Cross-Platform Computer Use Agents

Oct 22, 2025

Existing computer-use agents rely on platform-specific interfaces, hindering cross-environment deployment. To address this, we propose the first general-purpose vision-driven cross-platform operating agent. Our method employs hierarchical context management, decoupled planning and execution, self-verification feedback, and a multi-attempt decision mechanism—enabling end-to-end cross-platform adaptation without task-specific fine-tuning. The agent uniformly supports web (WebVoyager/WebArena), desktop (OSWorld), and mobile (AndroidWorld) environments. It achieves state-of-the-art accuracy of 97.1%, 69.6%, 60.1%, and 87.1% across four benchmark datasets, respectively; under multi-attempt evaluation, its overall performance surpasses human baselines. This work establishes the first SOTA-level cross-platform generalization capability and enables robust long-horizon task adaptation and autonomous recovery—marking a significant advance in universal, vision-based agent design.

0 citationsRead paper

Surfer-H Meets Holo1: Cost-Efficient Web Agent Powered by Open Weights

Jun 03, 2025

This work addresses the challenge of building low-cost, high-performance web agents. We propose Surfer-H, a novel framework, and Holo1, a family of open-source vision-language models (VLMs) specifically designed for web UI understanding and navigation. Holo1 integrates synthetic data generation, self-generated agentic training data, and precise UI element localization. Evaluated on WebVoyager, it achieves a state-of-the-art accuracy of 92.2%; it also significantly outperforms prior methods on our newly introduced web interaction benchmark, WebClick, and established general UI benchmarks. To our knowledge, this is the first open-source VLM family dedicated to web UI perception. We fully release Holo1’s model weights and the WebClick dataset, achieving a Pareto-optimal trade-off between accuracy and inference cost. The framework provides an efficient, reproducible, end-to-end baseline for automated web task execution.

0 citationsRead paper
Recent publications

Latest Papers

DragOn: A Benchmark and Dataset for Drag-Based GUI Interactions

Jun 04, 2026

This work addresses the limited performance of existing GUI agents in drag-based interactions—such as text highlighting and slider manipulation—which stems primarily from the absence of large-scale, high-quality datasets capturing such operations. To bridge this gap, the authors present DragOn, the first systematically constructed benchmark dataset specifically designed for drag interactions, encompassing four representative scenarios with 286,000 training screenshots, 3.5 million task instances, and 2,000 evaluation samples. They further introduce an end-to-end drag localization training paradigm conditioned on screen images and natural language instructions. Fine-tuning the Qwen vision-language model on DragOn substantially improves performance on drag-related tasks, demonstrating the dataset’s effectiveness in enhancing model generalization to real-world digital interactions and filling a critical data void in complex, continuous GUI manipulation.

0 citationsRead paper

Best-of-Q: Improving VLM agents with Q-function Action Ranking at Inference

Jan 30, 2026

Existing vision-language model (VLM) agents exhibit limited adaptability in dynamic environments such as web navigation and incur high costs when fine-tuned. This work proposes a training-free, inference-time optimization approach that decouples action generation from selection: the VLM is frozen to generate candidate actions, while a lightweight, offline-trained Q-function reranks these candidates. Notably, this is the first method to directly employ a Q-function for action selection during inference without updating the underlying policy, enabling immediate performance gains. Evaluated on the WebVoyager benchmark, the approach boosts the success rate of Qwen2.5-VL-7B from 38.8% to 55.7% and GPT-4.1 from 82.4% to 88.8%, significantly outperforming baseline methods.

0 citationsRead paper

Surfer 2: The Next Generation of Cross-Platform Computer Use Agents

Oct 22, 2025

Existing computer-use agents rely on platform-specific interfaces, hindering cross-environment deployment. To address this, we propose the first general-purpose vision-driven cross-platform operating agent. Our method employs hierarchical context management, decoupled planning and execution, self-verification feedback, and a multi-attempt decision mechanism—enabling end-to-end cross-platform adaptation without task-specific fine-tuning. The agent uniformly supports web (WebVoyager/WebArena), desktop (OSWorld), and mobile (AndroidWorld) environments. It achieves state-of-the-art accuracy of 97.1%, 69.6%, 60.1%, and 87.1% across four benchmark datasets, respectively; under multi-attempt evaluation, its overall performance surpasses human baselines. This work establishes the first SOTA-level cross-platform generalization capability and enables robust long-horizon task adaptation and autonomous recovery—marking a significant advance in universal, vision-based agent design.

0 citationsRead paper

Surfer-H Meets Holo1: Cost-Efficient Web Agent Powered by Open Weights

Jun 03, 2025

This work addresses the challenge of building low-cost, high-performance web agents. We propose Surfer-H, a novel framework, and Holo1, a family of open-source vision-language models (VLMs) specifically designed for web UI understanding and navigation. Holo1 integrates synthetic data generation, self-generated agentic training data, and precise UI element localization. Evaluated on WebVoyager, it achieves a state-of-the-art accuracy of 92.2%; it also significantly outperforms prior methods on our newly introduced web interaction benchmark, WebClick, and established general UI benchmarks. To our knowledge, this is the first open-source VLM family dedicated to web UI perception. We fully release Holo1’s model weights and the WebClick dataset, achieving a Pareto-optimal trade-off between accuracy and inference cost. The framework provides an efficient, reproducible, end-to-end baseline for automated web task execution.

0 citationsRead paper