An Agency-Transferring Model-Free Policy Enhancement Technique
This work addresses the inefficiency of existing reinforcement learning methods that often disregard available suboptimal baseline policies, resulting in high training costs and low task success rates. We propose a model-free policy augmentation framework that leverages a dynamic arbitration mechanism: during early training, control is delegated to a functional baseline policy to ensure goal reachability, and is gradually transferred to a learnable policy, ultimately yielding a high-performance policy independent of the baseline. We formally define functional baselines for the first time and integrate probabilistic reachability analysis to design the transfer mechanism, providing theoretical guarantees on the lower bound of goal achievement probability for the final policy. Experiments on continuous control benchmarks demonstrate that our method achieves competitive or superior returns compared to state-of-the-art approaches while consistently maintaining the highest goal success rate throughout both training and standalone deployment.