psRL: Efficient Training for Agentic AI via Training-Time Prefix Sharing

📅 2026-08-26
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文针对现代代理AI训练中更新阶段成为瓶颈的问题,提出psRL系统,通过训练样本前缀共享优化工作负载调度和内存管理,提高训练效率。
📝 Abstract
In modern agentic AI training, the system bottleneck is shifting from rollout to update. Emerging sampling strategies such as tree-structured and step-wise RL greatly increase training sample volume while incurring relatively low marginal rollout cost, causing the update phase to dominate the end-to-end execution time. Crucially, this shift exposes a new optimization opportunity, as production traces reveal substantial prefix redundancy across training samples. In this paper, we propose psRL (prefix sharing for RL), a new training system for agentic AI designed to exploit prefix redundancy among training samples. Leveraging the global visibility and data immutability inherent to the update phase, psRL achieves efficient workload scheduling and memory management for distributed training. Specifically, psRL introduces two novel prefix-sharing mechanisms that enable flexible, fine-grained workload distribution across GPU workers, simultaneously optimizing prefix reuse and achieving load balancing. Moreover, psRL implements a new underlying KV cache manager that facilitates adaptable block-size allocation and dynamic KV caching, maximizing memory utilization while maintaining a high prefix hit rate. Evaluations using production traces demonstrate that psRL outperforms existing systems by up to 5.2x in throughput. The source code will be publicly available soon.
Problem

Research questions and friction points this paper is trying to address.

agentic AI
prefix redundancy
training sample volume
update phase
workload scheduling
Innovation

Methods, ideas, or system contributions that make the work stand out.

prefix sharing
workload distribution
KV cache manager
memory utilization
M
Mianjie Yu
University of Macau
Z
Zizhao Mo
University of Macau
H
Huanyu Qu
University of Macau
Z
Zhirong Qian
University of Macau
H
Huanle Xu
University of Macau
C
Cen Li
Independent Researchers
Zifeng Zhao
Zifeng Zhao
University of Notre Dame
change-point analysisonline learningcopulaextreme value theorytime series
Z
Zhi Zhou
Independent Researchers
J
Jinhua Zhou
Independent Researchers
J
Jun Xie
Independent Researchers
C
Chengzhong Xu
University of Macau