Institution profile

China Merchants Research Institute of Advanced Technology

Academic institutionasia · cn
Official website
Research library2linked papers
Opportunities0open roles
Selected work

Representative Papers

Orcust: Stepwise-Feedback Reinforcement Learning for GUI Agent

Sep 22, 2025

Existing GUI agents face two critical challenges: unreliable reward signals and limited capacity for online trajectory generation—leading to unreliable reasoning and low data efficiency. To address these, we propose Orcust, a novel framework introducing two key mechanisms: (1) Principle-Constrained Reward Modeling (PCRM), which incorporates interpretable, human-aligned rules into the reinforcement learning reward function to ensure policy adherence to domain principles; and (2) Online Virtual Machine-Grounded Trajectory Construction (OVTC), which leverages a lightweight virtual machine to autonomously generate high-quality, structured interaction trajectories enabling fine-grained, stepwise training. Integrating chain-of-thought reasoning with LLM-driven rule feedback, Orcust achieves +22.2% and +23.9% improvements on ScreenSpot and ScreenSpot-Pro benchmarks, respectively. The framework significantly enhances reasoning reliability, task adaptability, and system scalability.

0 citationsRead paper

L0: Reinforcement Learning to Become General Agents

Jun 30, 2025

To address scalability and training efficiency bottlenecks in deploying large language models (LLMs) as autonomous agents for multi-turn, long-horizon tasks, this paper introduces L0—a fully end-to-end reinforcement learning framework. Methodologically, L0 features: (1) NB-Agent, an agent architecture adopting a “code-as-action” REPL execution paradigm; (2) Reinforcement Learning with Verifiable Rewards (RLVR), which directly elicits problem-solving capabilities from base models without supervised fine-tuning; and (3) a lightweight sandboxed concurrent agent pool enabling high-throughput, low-cost environment interaction. Evaluated on Qwen2.5-7B-Instruct, L0 achieves substantial improvements: SimpleQA accuracy rises from 30% to 80%, and HotpotQA from 22% to 41%. The framework is open-sourced, establishing a novel paradigm for scalable, efficient training of LLM-based autonomous agents.

0 citationsRead paper
Recent publications

Latest Papers

Orcust: Stepwise-Feedback Reinforcement Learning for GUI Agent

Sep 22, 2025

Existing GUI agents face two critical challenges: unreliable reward signals and limited capacity for online trajectory generation—leading to unreliable reasoning and low data efficiency. To address these, we propose Orcust, a novel framework introducing two key mechanisms: (1) Principle-Constrained Reward Modeling (PCRM), which incorporates interpretable, human-aligned rules into the reinforcement learning reward function to ensure policy adherence to domain principles; and (2) Online Virtual Machine-Grounded Trajectory Construction (OVTC), which leverages a lightweight virtual machine to autonomously generate high-quality, structured interaction trajectories enabling fine-grained, stepwise training. Integrating chain-of-thought reasoning with LLM-driven rule feedback, Orcust achieves +22.2% and +23.9% improvements on ScreenSpot and ScreenSpot-Pro benchmarks, respectively. The framework significantly enhances reasoning reliability, task adaptability, and system scalability.

0 citationsRead paper

L0: Reinforcement Learning to Become General Agents

Jun 30, 2025

To address scalability and training efficiency bottlenecks in deploying large language models (LLMs) as autonomous agents for multi-turn, long-horizon tasks, this paper introduces L0—a fully end-to-end reinforcement learning framework. Methodologically, L0 features: (1) NB-Agent, an agent architecture adopting a “code-as-action” REPL execution paradigm; (2) Reinforcement Learning with Verifiable Rewards (RLVR), which directly elicits problem-solving capabilities from base models without supervised fine-tuning; and (3) a lightweight sandboxed concurrent agent pool enabling high-throughput, low-cost environment interaction. Evaluated on Qwen2.5-7B-Instruct, L0 achieves substantial improvements: SimpleQA accuracy rises from 30% to 80%, and HotpotQA from 22% to 41%. The framework is open-sourced, establishing a novel paradigm for scalable, efficient training of LLM-based autonomous agents.

0 citationsRead paper