Institution profile

Reliant AI

Industry researchnorthamerica · us
Official website
Research library1linked papers
Opportunities0open roles
Selected work

Representative Papers

Tapered Off-Policy REINFORCE: Stable and efficient reinforcement learning for LLMs

Mar 18, 2025

To address slow training speed, poor stability, reliance on KL regularization, and inefficient negative-sample utilization in reinforcement learning fine-tuning of large language models (LLMs), this paper proposes Tapered Off-Policy REINFORCE (TOPR). TOPR employs asymmetric tapered importance sampling to enable fully offline, unified modeling of positive and negative samples while eliminating KL regularization entirely. Its key contributions are: (i) the first tapered importance sampling mechanism; (ii) theoretical and empirical identification of the implicit policy-distribution regularization effect induced by the REINFORCE baseline under negative sampling; and (iii) the first effective integration of negative samples in offline RL, yielding substantial gains in validation accuracy. On GSM8K and MATH benchmarks, an 8B model trained with TOPR matches the performance of a KL-regularized 70B model, achieves significantly higher inference accuracy, improves data efficiency, eliminates “inference waste” from negative samples, and supports multi-round iterative refinement with consistent gains.

0 citationsRead paper
Recent publications

Latest Papers

Tapered Off-Policy REINFORCE: Stable and efficient reinforcement learning for LLMs

Mar 18, 2025

To address slow training speed, poor stability, reliance on KL regularization, and inefficient negative-sample utilization in reinforcement learning fine-tuning of large language models (LLMs), this paper proposes Tapered Off-Policy REINFORCE (TOPR). TOPR employs asymmetric tapered importance sampling to enable fully offline, unified modeling of positive and negative samples while eliminating KL regularization entirely. Its key contributions are: (i) the first tapered importance sampling mechanism; (ii) theoretical and empirical identification of the implicit policy-distribution regularization effect induced by the REINFORCE baseline under negative sampling; and (iii) the first effective integration of negative samples in offline RL, yielding substantial gains in validation accuracy. On GSM8K and MATH benchmarks, an 8B model trained with TOPR matches the performance of a KL-regularized 70B model, achieves significantly higher inference accuracy, improves data efficiency, eliminates “inference waste” from negative samples, and supports multi-round iterative refinement with consistent gains.

0 citationsRead paper