MI-Distillation: Selecting from Model-Interpolated Instruct-Reasoning Data Spectrum for Chain-of-Thought Distillation

📅 2026-08-30
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文通过提出MI-Distillation方法,解决了长链思维轨迹难以有效蒸馏到小模型的问题,该方法利用模型插值构建推理数据谱,并选择适合学生模型的学习路径。
📝 Abstract
Recent advances in large reasoning models (LRMs) have shown strong performance on complex problems through long chain-of-thought (Long CoT) reasoning. However, distilling such trajectories into smaller student models remains challenging: direct Long CoT supervision often provides limited gains and can be less effective than concise Short CoT rationales. In this work, we investigate this phenomenon from a gradient-centric perspective. Our analysis shows that Long CoT induces larger gradient magnitudes and more concentrated update directions than Short CoT, with this effect becoming more pronounced as student model capacity increases. These findings suggest that effective Long CoT distillation requires balancing the reasoning information density of reasoning trajectories with their distributional alignment to the student model. Motivated by this insight, we propose \textbf{M}odel \textbf{I}nterporlation \textbf{Distillation} (\textbf{MI-Distillation}), a framework that constructs a continuous Instruct-Reasoning data spectrum through model interpolation. To select suitable trajectories from this spectrum, we further introduce \textbf{Seq}uential \textbf{L}earnable \textbf{S}urprisal \textbf{S}core (\textbf{SeqLSS}), which favors reasoning paths that are both informative and learnable for the student. Extensive experiments on reasoning benchmarks show that MI-Distillation consistently improves small model CoT distillation over strong Long CoT baselines.
Problem

Research questions and friction points this paper is trying to address.

Long CoT
distillation
student models
gradient magnitudes
reasoning information density
Innovation

Methods, ideas, or system contributions that make the work stand out.

Model-Interpolation
Chain-of-Thought Distillation
Instruct-Reasoning Data Spectrum
Sequential Learnable Surprisal Score
🔎 Similar Papers
No similar papers found.
Y
Yangsong Lan
College of Artificial Intelligence, Nanjing University of Aeronautics and Astronautics, Nanjing, China; The Key Laboratory of Brain-Machine Intelligence Technology, Ministry of Education, Nanjing, China
R
Renkai Hu
College of Artificial Intelligence, Nanjing University of Aeronautics and Astronautics, Nanjing, China; The Key Laboratory of Brain-Machine Intelligence Technology, Ministry of Education, Nanjing, China
H
HongKai Zheng
College of Artificial Intelligence, Nanjing University of Aeronautics and Astronautics, Nanjing, China; The Key Laboratory of Brain-Machine Intelligence Technology, Ministry of Education, Nanjing, China
B
Bo Zhang
College of Artificial Intelligence, Nanjing University of Aeronautics and Astronautics, Nanjing, China; The Key Laboratory of Brain-Machine Intelligence Technology, Ministry of Education, Nanjing, China
R
Renzhi Wang
College of Artificial Intelligence, Nanjing University of Aeronautics and Astronautics, Nanjing, China; The Key Laboratory of Brain-Machine Intelligence Technology, Ministry of Education, Nanjing, China
Hongliang Dai
Hongliang Dai
Nanjing University of Aeronautics and Astronautics
Information ExtractionLLMsKnowledge Graph
P
Piji Li
College of Artificial Intelligence, Nanjing University of Aeronautics and Astronautics, Nanjing, China; The Key Laboratory of Brain-Machine Intelligence Technology, Ministry of Education, Nanjing, China