Online Draft Co-Training for Speculative Decoding in Large-Scale, Long-Context RL Post-Training

📅 2026-09-07
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文解决了大规模长上下文模型中投机解码在线协同训练的问题,通过改进因果上下文并行和管道并行方法提高效率与准确性。
📝 Abstract
Speculative decoding accelerates rollout generation, which dominates the cost of reinforcement learning (RL) post-training. Online co-training can further increase the draft's accuracy, yielding greater speedups. However, scaling this approach to co-training on large models with long contexts poses two obstacles: (1) branch attention is unsupported by standard causal context-parallel (CP) implementations, and (2) target features span across pipeline-parallel (PP) stages. We address both with an end-to-end system for large-scale online draft co-training. For CP, we extend packed, load-balanced zigzag ring attention by merging rank-local branch attention with causal main-sequence attention. For PP, TapChannel transports intermediate target features across stages via a separate path, leaving the pipeline schedule unaffected. Experiments demonstrate that co-trained drafts closely track the policy baseline while delivering substantial rollout and end-to-end speedups across model scales up to 122B. Our CP design achieves strong scaling at 256K tokens with significant memory savings over prior work, and our PP transport incurs modest overhead. Code can be found at https://github.com/NVIDIA-NeMo/RL/issues/3698.
Problem

Research questions and friction points this paper is trying to address.

speculative decoding
reinforcement learning post-training
online co-training
large models
long contexts
Innovation

Methods, ideas, or system contributions that make the work stand out.

speculative decoding
online co-training
causal context-parallel
pipeline-parallel
TapChannel
🔎 Similar Papers
No similar papers found.