SimpleOPD: Simple Tokenizer-Agnostic On-Policy Distillation for Long-Context Reasoning

📅 2026-08-14
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of reasoning distillation between long-context teachers and short-context students arising from tokenizer mismatches and distribution shifts. We propose a tokenizer-agnostic on-policy distillation framework that incorporates shared text-space alignment, reference KL-divergence loss, and an end-of-sequence advantage masking mechanism to effectively mitigate length explosion and training instability. Experimental results demonstrate that this approach significantly enhances mathematical proving capabilities in short-context models; notably, Intern-S2 achieves a 21.2-point gain on ProofBench, surpassing Gemini-2.5-Pro, while yielding consistent improvements across scientific reasoning benchmarks.
📝 Abstract
On-policy distillation (OPD) offers a promising way to transfer reasoning capabilities from stronger teacher models, but applying it to long-context reasoning teachers and short-context students introduces practical challenges, including tokenizer mismatch, teacher-student distribution mismatch, response length explosion, and training instability. In this work, we study this setting by transferring proof-reasoning capabilities from the long-context reasoning model SU-01 to short-context student models. To handle tokenizer differences, we perform OPD in a shared text space and align only tokens that occupy identical text spans under the student and teacher tokenizers. To mitigate the problem of excessive generation length and frequent truncation, we introduce a student reference KL loss and mask the advantages of special termination tokens such as </think> and <|im_end|>. This strategy constrains the student from drifting excessively from its initial policy, thereby mitigating the teacher-student distribution mismatch problem and fostering steady length growth. Experiments on both same-family and different-family student models, including Qwen3, Qwen3.5, Intern-S2, GLM-4.7, Gemma-4, show consistent gains in mathematical reasoning, especially natural-language math proving. Notably, Intern-S2-Preview improves by 21.2 points on ProofBench, reaching 55.2 and surpassing Gemini-2.5-Pro. It also improves on science benchmarks such as HLE and HiPhO, suggesting that OPD transfers reasoning capabilities that generalize beyond the mathematical training domain.
Problem

Research questions and friction points this paper is trying to address.

On-Policy Distillation
Long-Context Reasoning
Tokenizer Mismatch
Distribution Mismatch
Training Instability
Innovation

Methods, ideas, or system contributions that make the work stand out.

On-Policy Distillation
Tokenizer-Agnostic
Long-Context Reasoning
Student Reference KL Loss
Proof Reasoning
🔎 Similar Papers
No similar papers found.
H
Haonan He
SU-01 Team, Shanghai Artificial Intelligence Laboratory
H
Haodi Lei
SU-01 Team, Shanghai Artificial Intelligence Laboratory
Yun Luo
Yun Luo
Shanghai AI Lab
natural language processinggraph neural network
H
Haoran Zhang
SU-01 Team, Shanghai Artificial Intelligence Laboratory
S
Shunkai Zhang
SU-01 Team, Shanghai Artificial Intelligence Laboratory
Yizhuo Li
Yizhuo Li
The University of Hong Kong
Shengji Tang
Shengji Tang
CUHK & Fudan University & Shanghai AI Lab
machine learningmodel compressionmodel design
Z
Zhilin Wang
SU-01 Team, Shanghai Artificial Intelligence Laboratory
Runzhe Zhan
Runzhe Zhan
Ph.D. Candidate, University of Macau
Machine TranslationLanguage ModelsMultilinguality
Lei Bai
Lei Bai
Shanghai AI Laboratory
Foundation ModelScience IntelligenceMulti-Agent SystemAutonomous Discovery
Ganqu Cui
Ganqu Cui
Shanghai AI Lab
LLM AlignmentReinforcement Learning
Fangchen Yu
Fangchen Yu
Ph.D Candidate, The Chinese University of Hong Kong, Shenzhen
Satistical Machine LearningOptimizationAI for ScienceMLLM
Yafu Li
Yafu Li
The Chinese University of Hong Kong
ReasoningTrustworthy AIMultilinguality
P
Peng Ye
SU-01 Team, Shanghai Artificial Intelligence Laboratory
N
Ning Ding
SU-01 Team, Shanghai Artificial Intelligence Laboratory
Y
Yu Cheng
SU-01 Team, Shanghai Artificial Intelligence Laboratory