Rethinking Reverse KL as Adaptive Entropy Distillation

📅 2026-08-05
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究解决了知识蒸馏中模仿准确性和生成鲁棒性的平衡问题,通过重新审视并改进RKL方法,提出自适应熵蒸馏(AED),动态调整学生模型的模仿强度。
📝 Abstract
Knowledge distillation (KD) is widely used to transfer the capabilities of large language models (LLMs) to smaller students, but existing objectives often struggle to balance faithful imitation and robust generation. In particular, existing methods mainly combine FKL and RKL, overlooking that RKL itself provides a mechanism for adjusting the student's imitation strength. Motivated by this, we revisit on-policy Reverse Kullback-Leibler (RKL) distillation and decompose its objective into a teacher-fitting term and a student-entropy term, without introducing an explicit FKL branch. We show theoretically that the token-level optimal student distribution corresponds to a tempered variant of the teacher distribution, where the adaptive weight controls the trade-off between mode-seeking and uncertainty preservation. Guided by this insight, we propose \textbf{Adaptive Entropy Distillation (AED)}, which uses the teacher's entropy to dynamically calibrate token-level imitation strength. Experiments on instruction-following and mathematical reasoning benchmarks demonstrate that AED achieves superior overall performance and generally improves teacher--student distributional and entropy alignment.
Problem

Research questions and friction points this paper is trying to address.

Knowledge Distillation
Reverse KL
Adaptive Entropy
Innovation

Methods, ideas, or system contributions that make the work stand out.

Adaptive Entropy Distillation
RKL
Entropy Calibration
Knowledge Distillation
🔎 Similar Papers
No similar papers found.
S
Shizhen Li
School of Computer Science and Engineering, Sun Yat-sen University, Guangzhou, China
Z
Zhiyu Shen
School of Computer Science and Engineering, Sun Yat-sen University, Guangzhou, China
Y
Yuyin Lu
School of Computer Science and Engineering, Sun Yat-sen University, Guangzhou, China
Y
Yunhe Pang
School of Computer Science and Engineering, Sun Yat-sen University, Guangzhou, China
J
Jielin Song
School of Computer Science and Engineering, Sun Yat-sen University, Guangzhou, China
Yanghui Rao
Yanghui Rao
Sun Yat-sen University
Text MiningTopic ModelingRepresentation Learning
Fu Lee Wang
Fu Lee Wang
Hong Kong Metropolitan University
AIData ScienceLearning Technology