Vanishing L2 regularization for the softmax Multi Armed Bandit
This work addresses the long-standing challenge of characterizing the convergence behavior of L2-regularized softmax policy gradient methods as the regularization coefficient vanishes. Focusing on the multi-armed bandit setting, the paper establishes the first rigorous convergence analysis framework for softmax policy gradient under vanishing L2 regularization, circumventing the conventional reliance on convexity assumptions. Through a combination of theoretical analysis and numerical experiments, the study formally proves that the algorithm converges even as the regularization parameter approaches zero. Furthermore, extensive evaluations on standard benchmarks demonstrate that this vanishing-regularization regime consistently outperforms both unregularized and fixed-regularization strategies, highlighting its practical efficacy and theoretical significance.