Harness-RL: Black-Box Reinforcement Learning with Action-Args Decoupling for Central-Agent Multi-Agent Harnesses

📅 2026-08-30
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出Harness-RL框架,通过冲突感知策略优化与黑盒轨迹构建解决多智能体系统中中心智能体的训练问题,特别是在动作和参数解耦及动态调度上。
📝 Abstract
Large language model agents increasingly solve long-horizon tasks through multi-agent harnesses in which a central agent coordinates specialized sub-agents, tools, and environments. Training the central policy in such a harness raises two challenges. First, an action label is a low-cardinality decision, whereas its args form a high-dimensional conditional sequence; optimizing both with a shared sequence-level signal can produce conflicting gradients. Second, dynamic scheduling creates interdependent sessions with branches, parallel calls, and rewritten contexts, which cannot be faithfully reduced to one flat token sequence. We introduce Harness-RL, a structured reinforcement learning framework that combines Conflict-Aware Policy Optimization (CAPO) with interface-level black-box trajectory construction. The black-box component captures Interface Call Records, builds per-session prefix trees, and aligns outcome and process rewards with trainable tokens. CAPO uses forward activations to identify parameter partitions associated with action and args tokens, then routes their policy gradients to the corresponding subspaces. Harness-RL supports both central-only and joint multi-agent training. Across seven multi-hop question answering and agentic retrieval benchmarks, it reaches average F1 scores of 42.93 and 47.79 with Qwen2.5-1.5B and Qwen2.5-3B, respectively, while ablations validate the contribution of CAPO and favor central-only optimization in the evaluated setting. Our code is available at https://github.com/jiangxinke/Harness-RL.
Problem

Research questions and friction points this paper is trying to address.

Reinforcement Learning
Multi-Agent Systems
Central Agent
Gradient Conflict
Dynamic Scheduling
Innovation

Methods, ideas, or system contributions that make the work stand out.

Conflict-Aware Policy Optimization
Black-Box Trajectory Construction
Interface Call Records
Prefix Trees
🔎 Similar Papers
No similar papers found.
X
Xinke Jiang
National Engineering Research Center of Software Engineering, Peking University, Beijing, China
Zhixin Zhang
Zhixin Zhang
Ph.D of Robotics, University of Manchester
SLAMVINSLIOSensor FusionRobotics
Z
Zhibang Yang
National Engineering Research Center of Software Engineering, Peking University, Beijing, China
J
Jiaran Gao
National Engineering Research Center of Software Engineering, Peking University, Beijing, China
R
Rihong Qiu
National Engineering Research Center of Software Engineering, Peking University, Beijing, China
S
Shijin Chen
Guangxi Land and Resources Planning and Design Group Co., Ltd, Guangxi, China
Xu Chu
Xu Chu
Peking University
Machine learningData mining
Junfeng Zhao
Junfeng Zhao
Assistant Professor at Arizona State University, Director of BELIV Lab
Connected & Automated VehicleMotion Planning & ControlsElectric VehiclesAI/ML
Y
Yasha Wang
Peking University Information Technology Institute (Tianjin Binhai), Tianjin, China